System

The system addresses data integration and analysis challenges by preprocessing, training a multimodal AI model, and generating real-time alerts to enhance safety and efficiency in factories and construction sites.

JP2026022270APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123787
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing systems fail to efficiently integrate and analyze diverse data from on-site sensors and historical records to optimize safety and efficiency in factories and construction sites, leading to inefficiencies and safety risks.

Method used

A system that collects data from multiple input devices, preprocesses it, integrates it into a unified format, trains a multimodal AI model, and generates real-time analysis and alerts to improve efficiency and safety.

Benefits of technology

The system significantly enhances work efficiency and safety by identifying inefficiencies and dangers promptly, providing actionable suggestions, and improving accuracy through user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022270000001_ABST
    Figure 2026022270000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting information from an input device installed in a field; means for preprocessing the collected information; means for integrating the preprocessed information to generate a data set; means for learning a model using the data set; means for analyzing real-time data in the field to determine efficiency and safety; and means for generating and notifying an improvement proposal and an emergency alert based on a result of the determination.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Achieving both efficiency and safety in on-site work is extremely important in the operation of factories and construction sites. However, these sites are rife with risks and inefficiencies, and many sites in particular do not fully utilize data from past accidents and near-miss incidents. As a result, safety is threatened and work efficiency declines. The present invention aims to optimize efficiency and safety at on-site sites, detect inefficient work and dangerous situations in advance, and propose improvements. [Means for solving the problem]

[0005] The present invention is a system including: means for collecting information from input devices installed on-site; means for preprocessing the collected information; means for integrating the preprocessed information to generate a dataset; means for training a model using the dataset; means for analyzing real-time data from the site and determining efficiency and safety; and means for generating and notifying improvement proposals and emergency alerts based on the determination results. This system collects and analyzes past accident data, near-miss event logs, work procedures, completed volume data, and machine specification data, and analyzes the integrated data using multimodal AI to achieve an optimal balance between efficiency and safety.

[0006] An "input device" is a device installed to collect data such as video, audio, temperature, and vibration from the site.

[0007] "Preprocessing" refers to the process of removing noise and converting formats from collected data.

[0008] A "dataset" is a collection of pre-processed data that has been integrated into a single unified format.

[0009] A "model" is an algorithm that uses machine learning and deep learning techniques to analyze data and learn specific patterns.

[0010] "Real-time data" refers to the latest data collected from the field and is information that is analyzed without time delay.

[0011] "Efficiency" is the degree to which a task or process is carried out using fewer resources (time, manpower, energy, etc.).

[0012] "Safety" refers to work and environment conditions that minimize the risk of accidents, injuries, and breakdowns.

[0013] "Analysis" is the act of examining collected data and extracting meaningful information.

[0014] "Improvement proposals" are specific action plans based on the analysis results to improve work efficiency and safety at the site.

[0015] An "emergency alert" is a warning message that quickly notifies you when a dangerous situation occurs on-site.

[0016] "Past accident data" refers to recorded information about accidents that have previously occurred at the site.

[0017] A "near miss incident log" is a record of potentially dangerous incidents that did not actually result in an accident.

[0018] A "work procedure manual" is a document that shows the specific steps for performing a specific task.

[0019] "Output data" is data that indicates the results of production or work at a specific time or period.

[0020] "Machine specification data" refers to technical documentation regarding the performance, functions, settings, etc. of machines and equipment.

[0021] "Multimodal AI" is an artificial intelligence technology that can integrate and analyze data in different formats, such as text, images, and audio. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0024] First, the terms used in the following description will be explained.

[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0030] [First embodiment]

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0043] This invention describes an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement proposals and emergency alerts. The operation of this system's program is explained below in natural language.

[0044] Data Collection Overview

[0045] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[0046] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[0047] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0048] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[0049] Data Preprocessing Overview

[0050] 3. The server performs preprocessing on the collected data, specifically removing noise from the audio data, deleting unnecessary frames from the video data, and analyzing the text data.

[0051] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[0052] 4. The server converts the different types of data into a unified format.

[0053] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[0054] Overview of Data Integration and Model Training

[0055] 5. The server aggregates the pre-processed data to create a single aggregated dataset.

[0056] Example: The server stores cleaned video data, audio data, sensor data, and text data as a single dataset.

[0057] 6. The server trains the multimodal AI model using the integrated dataset.

[0058] Example: The server uses a deep learning framework to train models and improve its efficiency and safety decision-making capabilities.

[0059] Real-time analysis and improvement suggestions overview

[0060] 7. The server analyzes the collected real-time data and monitors the situation on site.

[0061] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0062] 8. Based on the analysis results, the server identifies inefficient areas and potential dangers and generates appropriate improvement proposals.

[0063] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[0064] 9. The server will immediately alert the user if an imminent danger occurs.

[0065] Example: The server detects signs of a fire from video data and immediately sends a warning message to the administrator's terminal.

[0066] Notifications and Feedback Overview

[0067] 10. The server notifies the user of the generated improvement suggestions and emergency alerts.

[0068] Example: The server sends a message to the user's device saying, "Introducing a new work procedure will improve efficiency."

[0069] 11. The user checks the notified suggestions and alerts and takes the necessary action.

[0070] Example: A user reviews proposed work procedure changes and instructs field workers on the new procedures.

[0071] 12. Users provide feedback to the server on suggestions and alerts, helping to improve the system's accuracy.

[0072] Example: A user fills in a feedback form with a newly encountered problem and submits it to the server.

[0073] The above is a specific embodiment of the system of the present invention, which can significantly improve efficiency and safety in factories and construction sites.

[0074] The processing flow will be explained below.

[0075] Step 1:

[0076] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[0077] Step 2:

[0078] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[0079] Step 3:

[0080] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements. For example, it deletes frames in which human movement cannot be confirmed.

[0081] Step 4:

[0082] The server analyzes text data (accident reports, near-miss incident logs, work procedure manuals, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[0083] Step 5:

[0084] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0085] Step 6:

[0086] The server then combines the pre-processed data to create a single integrated dataset, for example audio, video, temperature, and vibration data, which is stored in a cloud database and prepared for analysis.

[0087] Step 7:

[0088] The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0089] Step 8:

[0090] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal conditions from real-time video and extracting abnormal sounds from audio data.

[0091] Step 9:

[0092] The server uses the analysis results to identify areas of inefficiency and potential dangers. For example, if a location where workers frequently stop or excessive vibration is detected, it will identify that as a potential danger.

[0093] Step 10:

[0094] The server generates improvement suggestions aimed at improving efficiency and safety, for example by changing the layout of machines or proposing new work procedures.

[0095] Step 11:

[0096] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[0097] Step 12:

[0098] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[0099] Step 13:

[0100] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0101] Example 1

[0102] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0103] In factories and construction sites, real-time data collection and analysis is necessary to ensure both work efficiency and safety, but existing systems are difficult to adequately address this. There is also a need to integrate data from a variety of sensors and past data, and quickly generate appropriate improvement proposals and emergency alerts. Conventional technologies face challenges in unifying different data formats, removing noise, and improving the accuracy of real-time analysis.

[0104] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0105] In this invention, the server includes means for collecting data from input devices installed on-site, preprocessing means for removing noise and unnecessary data from the collected data, means for converting data of different formats into a unified format, means for integrating the preprocessed data to generate a dataset, means for training a multimodal model using the dataset, means for analyzing real-time data to determine efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting feedback from users to improve the accuracy of the system, thereby enabling significant improvements in work efficiency and safety at factories and construction sites.

[0106] A "server" is a central processing system that collects, processes, and analyzes data from various input devices installed in factories and construction sites, and notifies and presents the results to users.

[0107] "Input devices" are devices installed on-site to collect data, and include 3D cameras, microphones, temperature sensors, vibration sensors, etc.

[0108] The "data collection means" is a mechanism for continuously acquiring data from the input device and transmitting it to the server.

[0109] "Noise reduction means" refers to the algorithms and filtering processes used to remove background noise from collected audio data.

[0110] "Means for removing unnecessary data" refers to steps for eliminating frames in which no specific movement is detected or meaningless information from the video data.

[0111] "Preprocessing means" refers to a technique that includes processes such as noise removal and deletion of unnecessary data on collected raw data to make it ready for analysis.

[0112] "Means of converting into a unified format" refers to technology that unifies data collected in different formats into a consistent format, enabling integrated analysis.

[0113] "Dataset generation means" refers to the process of integrating preprocessed data and generating and saving a single integrated dataset.

[0114] A "multimodal model" is an artificial intelligence model that simultaneously analyzes different types of data (e.g., video, audio, text, sensor information) and learns their interrelationships.

[0115] "Model training methods" are techniques for training AI models using preprocessed and integrated datasets to improve their accuracy.

[0116] "Real-time data analysis means" refers to technology that instantly analyzes collected data and evaluates and monitors the current on-site situation.

[0117] "Means for determining efficiency and safety" refers to an algorithm that evaluates on-site work efficiency and safety based on the results of analyzing real-time data and makes specific judgments.

[0118] The "improvement proposal generation means" is a process for generating specific proposals for improving work efficiency and safety based on the results of data analysis.

[0119] The "emergency alert notification means" is a mechanism for immediately issuing a warning to the user when a serious danger is detected.

[0120] The "feedback collection means" is a mechanism for collecting user opinions and suggestions for improvement on a server and using them to improve and enhance the accuracy of the system.

[0121] The present invention relates to a system for optimizing the efficiency and safety of work in factories and construction sites. This system is composed of a server, an input device, and a user terminal.

[0122] The server collects data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. This data includes video data, audio data, temperature data, and vibration data. For example, the 3D camera streams video of the site in real time, and the microphone captures audio from the site. The temperature and vibration sensors collect their respective environmental data.

[0123] The server removes noise from the collected data and deletes unnecessary data. This preprocessing improves the accuracy of the analysis. Specifically, a voice recognition algorithm is used to remove noise from the audio data, and a motion detection algorithm is used to delete unnecessary frames from the video data. In addition, PDF-formatted work procedures and past accident reports are converted into text format using OCR technology.

[0124] The preprocessed data is then converted into a unified format by the server. This conversion integrates data from different formats into a consistent format. The converted data is then stored on the server as a unified dataset and analyzed using a multimodal model.

[0125] A multimodal model is an artificial intelligence model that simultaneously analyzes different types of data, optimally analyzing different modalities (video, audio, text, and sensor information). The server uses this multimodal model to determine efficiency and safety. For example, real-time video analysis can detect worker movements and identify unnatural movements. Audio analysis can also detect abnormal sounds on-site.

[0126] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and proposes relocation of the machines. In addition, if an emergency alert is required, the server immediately sends a warning message to the user. For example, it detects signs of a fire from video data and immediately sends a warning message to the user's device.

[0127] Users check improvement suggestions and emergency alerts notified by the server. Based on the notified information, they can change work procedures or instruct on-site workers on emergency responses. Users can also provide feedback on suggestions and alerts to the server, contributing to improving the accuracy of the system. For example, users can enter newly discovered problems or areas for improvement into a feedback form and send it to the server.

[0128] As a result, the efficiency and safety of work at factories and construction sites can be significantly improved. The system of the present invention can quickly and accurately monitor the situation at the site by analyzing real-time data, and provide appropriate improvement suggestions and emergency alerts.

[0129] An example of a prompt is, "Please analyze the current factory data and generate suggestions to improve safety." Using this prompt, the generative AI model can make appropriate analyses and suggestions.

[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0131] Step 1:

[0132] The server collects data from input devices installed on-site. These include a 3D camera, microphone, temperature sensor, and vibration sensor. The server receives a real-time video stream from the 3D camera and simultaneously captures audio data from the microphone. Data from the temperature sensor and vibration sensor is also sent to the server. This allows for an immediate understanding of the situation on-site.

[0133] Input: 3D camera image data, microphone audio data, temperature sensor and vibration sensor data

[0134] Output: Collected raw data (video, audio, temperature, vibration)

[0135] Step 2:

[0136] The server performs preprocessing on the collected data. Specifically, it removes noise from the audio data and deletes unnecessary frames from the video data. To remove noise from the audio data, it uses a speech recognition algorithm to remove background noise. For the video data, it uses a motion detection algorithm to remove inactive frames. Furthermore, it converts work procedures and accident reports into text data using OCR.

[0137] Input: Collected raw data (video, audio, temperature, vibration), work procedures, accident reports

[0138] Output: Preprocessed data (noise-removed audio, video with unnecessary frames deleted, materials converted into text data)

[0139] Step 3:

[0140] The server converts the pre-processed data into a unified format, which aligns data from different formats into a consistent format, such as text data converted from PDF or sensor data converted to JSON format.

[0141] Input: Preprocessed data

[0142] Output: Dataset converted to a unified format

[0143] Step 4:

[0144] The server then combines the data in a unified format to generate a single comprehensive dataset. All data is integrated based on timestamps and stored in a database. This dataset is then used for model training and real-time analysis.

[0145] Input: Data converted into a unified format

[0146] Output: Unified dataset

[0147] Step 5:

[0148] The server trains a multimodal AI model using the integrated dataset. It uses a deep learning framework (e.g., TensorFlow) to analyze patterns in the data and improve its ability to make safety and efficiency decisions. This training allows the model to analyze different types of data simultaneously.

[0149] Input: Integrated dataset

[0150] Output: Trained multimodal AI model

[0151] Step 6:

[0152] The server analyzes the collected real-time data and monitors the situation on-site. Real-time video analysis detects worker movements and identifies unnatural movements. Audio analysis detects abnormal sounds.

[0153] Input: Real-time data (video, audio, temperature, vibration)

[0154] Output: Analysis results (on-site condition monitoring data)

[0155] Step 7:

[0156] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and suggests changing the machine's location. It also immediately sends a warning message to the user if an emergency alert is required. For example, it detects signs of a fire from video data and sends a warning message to the user's device.

[0157] Input: Analysis results

[0158] Output: Improvement suggestions and emergency alerts

[0159] Step 8:

[0160] The server notifies the generated improvement suggestions and emergency alerts to the user's device, which has the function to receive and display the suggestions and alerts.

[0161] Input: Improvement suggestions and emergency alerts

[0162] Output: Notification to user terminal

[0163] Step 9:

[0164] The user checks the notified improvement proposals and alerts and takes the necessary action. For example, the user checks the proposed changes to work procedures and instructs the on-site workers on the new procedures. In the case of an emergency alert, the user can quickly implement countermeasures.

[0165] Input: Improvement suggestions and emergency alerts

[0166] Output: User response

[0167] Step 10:

[0168] Users can provide feedback on suggestions and alerts to the server, helping to improve the system's accuracy. This feedback is reflected in the next model training.

[0169] Input: User feedback

[0170] Output: Feedback data to the server

[0171] The above is the specific processing flow of this system.

[0172] (Application example 1)

[0173] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0174] To maximize the efficiency and safety of on-site work, it is important to monitor the situation in real time and quickly and accurately identify inefficient or dangerous areas. However, conventional systems have difficulty integrating dispersed data and analyzing it in real time, resulting in low accuracy in improvement proposals and emergency alerts. Furthermore, a monitoring system using advanced data analysis technology was required to enable robots to work safely and effectively on-site.

[0175] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0176] In this invention, the server includes means for collecting information from input devices installed on-site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for training a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting, analyzing, and monitoring robot sensor data in real time. This makes it possible to optimize the efficiency and safety of on-site work, reducing risks and improving work safety, particularly at sites where robots work.

[0177] "Worksite" refers to a place where work is carried out, such as a factory or construction site.

[0178] "Input devices" refers to devices such as sensors and cameras installed on-site to collect data.

[0179] "Preprocessing" refers to the process of removing unnecessary parts from collected data and converting it into a format that is easy to analyze.

[0180] "Dataset" refers to a set of data that is collected, preprocessed, and integrated from multiple data sources.

[0181] A "model" is an algorithm that learns from collected and preprocessed data to accomplish a specific task.

[0182] "Real-time data" refers to data that is collected continuously along an ongoing timeline.

[0183] "Efficiency" refers to the degree to which work is carried out smoothly and without waste.

[0184] "Safety" refers to the degree to which work is carried out without hazards.

[0185] "Improvement proposals" refer to recommendations for improving efficiency or safety.

[0186] "Emergency Alert" means an emergency notification of immediate danger at a site.

[0187] "Robot" refers to a mechanical device that is programmed to perform tasks automatically.

[0188] "Sensor data" refers to data such as temperature, vibration, audio, and video collected by sensors.

[0189] "Monitoring" refers to the act of continuously observing the situation at a site and detecting any changes.

[0190] "Camera" refers to a device that captures images and collects them as digital data.

[0191] "Microphone" refers to a device for collecting sound.

[0192] "Temperature sensor" refers to a device that measures temperature and collects it as digital data.

[0193] A "vibration sensor" refers to a device that detects vibrations and collects them as digital data.

[0194] "Real-time" means happening at the present time, without delay.

[0195] "Notification" refers to the act of sending a message to convey information.

[0196] "Server" refers to a computer system for collecting, preprocessing, integrating, analyzing, and notifying data.

[0197] This invention is realized using an AI system to optimize the efficiency and safety of work on-site. The system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient areas and dangerous spots, and generates improvement suggestions and emergency alerts.

[0198] 1. Program Overview

[0199] The server collects data from input devices such as 3D cameras, microphones, temperature sensors, and vibration sensors. It then preprocesses and integrates the collected data to generate a dataset. This dataset is used to train an AI model that analyzes real-time on-site data. Based on the analysis results, it makes efficiency and safety judgments, generates improvement suggestions and emergency alerts, and notifies the user's device or the robot's display.

[0200] 2. Hardware and Software Used

[0201] 3D camera: Collects footage of the scene in real time.

[0202] Microphone: Collects audio data.

[0203] Temperature Sensor: Collects temperature data in the field.

[0204] Vibration Sensor: Collects vibration data.

[0205] Server: Collects, preprocesses, integrates, analyzes, trains models, and notifies data. In particular, it uses deep learning models using TensorFlow and Keras.

[0206] User device: Receives and displays improvement suggestions and emergency alerts.

[0207] Robot: Monitors on-site operations and presents the necessary information to users in an easy-to-understand format.

[0208] 3. Data processing and calculation

[0209] The server performs the following data processing.

[0210] Data collection: Collect data in real time from 3D cameras, microphones, temperature sensors, and vibration sensors.

[0211] Data preprocessing: Converting video data to grayscale, removing noise from audio data, detecting outliers in temperature and vibration data, etc.

[0212] Data integration: Integrate various pre-processed data to generate a single dataset.

[0213] Model training: Use TensorFlow and Keras to train a multimodal AI model based on the integrated dataset.

[0214] Real-time analysis: Trained models analyze collected real-time data to assess efficiency and safety.

[0215] Notification generation: Based on the analysis results, improvement suggestions and emergency alerts are automatically generated and sent to the user's device.

[0216] 4. Examples and prompts

[0217] A concrete example of this system in action is a robot in a factory. The robot operates safely by collecting data from 3D cameras and temperature sensors in the work area. The server analyzes this data in real time, identifies abnormal temperature rises or unnatural movements in the work area, and immediately sends an emergency alert to the manager's terminal.

[0218] Example prompt sentence:

[0219] "Please explain in detail how real-time data from 3D cameras and temperature sensors is analyzed to train AI models that determine efficiency and risk so that factory robots can continue to work safely."

[0220] By implementing this invention, the efficiency and safety of work in factories and construction sites can be significantly improved.

[0221] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0222] Step 1:

[0223] The server collects data from 3D cameras, microphones, temperature sensors, and vibration sensors installed on-site. Video data, audio data, temperature data, and vibration data obtained from input devices are input to the server. The collected data is in the form of a video stream for video, an audio stream for audio, and numerical data output from the sensors for temperature and vibration. This allows the real-time situation on-site to be accumulated on the server as digital data.

[0224] Step 2:

[0225] The server performs preprocessing on the collected data. This involves converting the video data to grayscale and removing noise from the audio data. It also detects and filters out abnormal values ​​in the temperature and vibration data. Specifically, OpenCV is used to convert each frame of the video data to grayscale, and an audio filtering algorithm is used to remove unwanted noise from the audio data. This results in clean data suitable for analysis.

[0226] Step 3:

[0227] The server then combines the pre-processed data into a single dataset, including video frames converted to grayscale, filtered audio data, and normalized temperature and vibration data, consolidating the various data types into a single, integrated format for further analysis.

[0228] Step 4:

[0229] The server trains an AI model using the integrated dataset. Specifically, it uses TensorFlow and Keras to train a deep learning model. The input data is a preprocessed and integrated dataset, and the output is a model for evaluating site efficiency and safety. This model learns based on historical data and real-time input data.

[0230] Step 5:

[0231] The server uses a trained AI model to analyze on-site data collected in real time. The input data is video data, audio data, temperature data, and vibration data collected in real time, and the model analyzes this data to evaluate efficiency and safety. The analysis results output a judgment that a particular task is inefficient or dangerous.

[0232] Step 6:

[0233] The server generates improvement suggestions and emergency alerts based on the analysis results. For example, if workers' movements in a specific area are unnatural, it will suggest changes to work procedures as an improvement suggestion. If temperature data is abnormally high, it will determine that there is a risk of fire and generate an emergency alert. These notifications are generated in a specific and actionable format.

[0234] Step 7:

[0235] The server notifies the generated improvement proposals and emergency alerts to the user's terminal or the robot's display. The notification content is displayed so that the user can immediately understand it and take measures. Specifically, the server sends email notifications using SMTP and displays warning messages on the robot's display.

[0236] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0237] This invention describes an AI system for optimizing the efficiency and safety of work in factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. Furthermore, this invention combines an emotion engine that recognizes the user's emotions, enabling it to provide appropriate feedback based on the user's emotional state.

[0238] Data Collection Overview

[0239] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. If there is a shortage, it also refers to backup data.

[0240] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[0241] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0242] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[0243] Data Preprocessing Overview

[0244] 3. The server performs noise reduction on the collected audio data, specifically filtering out background noise and extracting only the target audio.

[0245] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[0246] 4. The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[0247] 5. The server converts different data formats into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0248] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[0249] Overview of Data Integration and Model Training

[0250] 6. The server combines the pre-processed data to create a single combined data set, for example audio, video, temperature, and vibration data, and stores it in a cloud database, ready for analysis.

[0251] 7. The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0252] Example: The server uses past data to train an AI model to help determine future work efficiency and safety.

[0253] Introducing the Emotion Engine

[0254] 8. The server combines an emotion engine that recognizes the user's emotions and analyzes the user's emotional state. For example, it analyzes audio data acquired from a microphone and recognizes the user's emotional state (anger, sadness, joy, etc.).

[0255] Example: The server uses speech recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[0256] 9. The server analyzes the user's facial expressions from the video data and detects changes in emotions. For example, it analyzes video data acquired from a camera and determines whether the user is smiling or confused.

[0257] Example: The server uses a facial expression recognition algorithm to suggest appropriate actions to take if the user is feeling anxious.

[0258] Real-time analysis and improvement suggestions overview

[0259] 10. The server analyzes the collected real-time data and monitors the situation at the site. It detects worker movements and abnormal situations from real-time video and extracts abnormal sounds from audio data.

[0260] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0261] 11. Based on the analysis results, the server identifies inefficiencies and potential dangers and generates appropriate improvement proposals, such as changing the layout of machines or proposing new work procedures.

[0262] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[0263] 12. The server will immediately send an alert to the user if an emergency occurs. If it detects signs of a fire, it will immediately send a warning message to the administrator's terminal.

[0264] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[0265] Notifications and Feedback Overview

[0266] 13. The server notifies the user of the generated improvement suggestions and emergency alerts. The server adjusts the content and timing of the feedback based on the user's emotional state.

[0267] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[0268] 14. The user checks the suggestions and alerts sent from the server and takes necessary action. For example, the user communicates the proposed new work procedures to the field workers and instructs them to carry them out immediately.

[0269] Example: A user reviews the proposed new work procedure and instructs field workers on the new procedure.

[0270] 15. Users contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0271] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[0272] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[0273] The processing flow will be explained below.

[0274] Step 1:

[0275] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[0276] Step 2:

[0277] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[0278] Example: A server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone.

[0279] Step 3:

[0280] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements.

[0281] Example: The server removes still frames from video data and extracts only frames with movement.

[0282] Step 4:

[0283] The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases.

[0284] Example: The server analyzes accident reports and extracts important keywords such as "fall," "fire," and "machine malfunction."

[0285] Step 5:

[0286] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0287] Example: A server converts PDF files into text format to create a dataset for multimodal AI.

[0288] Step 6:

[0289] The server aggregates the pre-processed data to create a single aggregated data set.

[0290] Example: The server combines cleaned audio, video, sensor, and text data to create a unified dataset.

[0291] Step 7:

[0292] The server uses the combined data set to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0293] Example: The server uses past data to train AI models to improve future work efficiency and safety.

[0294] Step 8:

[0295] The server uses an emotion engine that recognizes the user's emotions to analyze the user's voice data and recognize the user's emotional state (anger, sadness, joy, etc.).

[0296] Example: A server uses voice recognition and emotion analysis algorithms to determine whether a user is stressed.

[0297] Step 9:

[0298] The server analyzes the user's facial expressions from the video data and detects changes in emotions.

[0299] Example: The server uses a facial expression recognition algorithm to determine whether the user is smiling, confused, etc.

[0300] Step 10:

[0301] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal situations from real-time video and extracting abnormal sounds from audio data.

[0302] Example: The server uses real-time video and audio analysis to detect worker movements and identify unnatural movements.

[0303] Step 11:

[0304] Based on the analysis results, the server identifies inefficient areas and potential danger points and generates appropriate improvement proposals.

[0305] Example: The server analyzes the locations where workers frequently stop and suggests relocating those locations.

[0306] Step 12:

[0307] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[0308] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[0309] Step 13:

[0310] The server notifies the user of the generated improvement suggestions and emergency alerts, and adjusts the content and timing of the feedback based on the user's emotional state.

[0311] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[0312] Step 14:

[0313] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[0314] Example: A user reviews a proposed new work procedure and communicates the new procedure to workers in the field.

[0315] Step 15:

[0316] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0317] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[0318] Example 2

[0319] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0320] Optimizing the efficiency and safety of on-site work requires real-time situational awareness and appropriate feedback. However, existing systems lack the ability to integrate and analyze collected data, making it difficult to respond quickly and accurately to specific issues. In addition, they are unable to provide feedback that takes into account the emotional state of workers, limiting further improvements in efficiency and safety.

[0321] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting information from input devices installed at the site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for learning a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for analyzing the emotional state of the user and providing appropriate feedback based on the emotional state. This makes it possible to analyze the situation at the site in real time, improve efficiency and safety, and provide appropriate feedback based on the emotional state of the worker.

[0322] "Worksite" refers to the location where work actually takes place, such as a factory or construction site.

[0323] "Input devices" refers to equipment including sensors, cameras, microphones, etc. that are installed on-site to collect information.

[0324] "Means for collecting information" refers to the method or technology for transmitting data obtained from the input device to the server.

[0325] "Preprocessing" refers to a series of steps taken to convert collected raw data into a form that is easier to analyze.

[0326] "Preprocessing means" refers to techniques and methods for performing preprocessing, such as noise removal, data transformation, and extraction of important information.

[0327] A "dataset" refers to a single set of data that is created by integrating multiple preprocessed data.

[0328] "Means for generating datasets" refers to methods and techniques for combining pre-processed data into a format that an AI model can learn from.

[0329] "Means of training the model" refers to the methods and techniques used to train the AI ​​model using the generated dataset.

[0330] "Real-time data" refers to the latest data collected from the field.

[0331] "Means of analysis" refers to the techniques and methods used to analyze collected data and determine its efficiency and safety.

[0332] "Decision result" refers to the conclusion or evaluation obtained based on the analyzed data.

[0333] "Improvement proposals" refer to specific proposals or methods for improving efficiency or safety.

[0334] An "urgent alert" is a warning issued when a danger or problem occurs that requires immediate action.

[0335] "Means for notification" refers to the methods and technologies for communicating generated improvement suggestions and emergency alerts to users.

[0336] "Emotional state" refers to a user's mental state or emotion, including, for example, joy, anger, sadness, and the like.

[0337] "Means for providing feedback" refers to methods and techniques for providing appropriate information or advice to a user based on the analyzed emotional state.

[0338] This invention is an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes the user's emotional state to provide appropriate feedback.

[0339] The system configuration is as follows:

[0340] Data collection

[0341] The server collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. For example, the server receives video streams from 3D cameras and simultaneously captures audio data from microphones. It also collects data on past accidents, near-miss incident logs, work procedures, completed volume data, and machine specification data.

[0342] Specific examples

[0343] The server simultaneously collects the video stream from the 3D camera and the audio data from the microphone.

[0344] The server retrieves past fall accident reports and work procedure manuals in PDF format and stores them in a database.

[0345] Data Preprocessing

[0346] The server performs noise reduction on the collected audio data. Specifically, it filters out background noise from the audio data and extracts only the target audio. It also removes frames from the video data in which no specific movement is detected.

[0347] In addition, the server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, and also converts data in different formats into a unified format.

[0348] Specific examples

[0349] The server uses a voice recognition algorithm to remove background noise from the voice data acquired from the microphone.

[0350] The server converts the PDF files into text format, generating a dataset that is easy to analyze.

[0351] Data integration and model training

[0352] The server combines the pre-processed data to create a single integrated data set, which is then used to train an AI model that uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0353] Specific examples

[0354] The server uses past data to train the AI ​​model, helping to determine future work efficiency and safety.

[0355] Introducing the Emotion Engine

[0356] The server combines an emotion engine to recognize the user's emotional state. For example, it analyzes audio data acquired from a microphone to recognize the user's emotional state (anger, sadness, joy, etc.). It also analyzes the user's facial expressions from video data to detect emotional fluctuations.

[0357] Specific examples

[0358] The server uses voice recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[0359] The server uses a facial expression recognition algorithm to suggest appropriate measures to address the user's anxiety.

[0360] Real-time analysis and improvement suggestions

[0361] The server analyzes the collected real-time data and monitors the situation at the site. Based on the analysis results, it identifies inefficiencies and potential dangers, generates improvement proposals, and immediately sends alerts if an emergency danger occurs, if necessary.

[0362] Specific examples

[0363] The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0364] If the server detects a fire, it will send a real-time alert to the administrator's smartphone.

[0365] Notifications and Feedback

[0366] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the feedback based on the user's emotional state.

[0367] The user checks the suggestions and alerts sent from the server and takes necessary action. The user also provides feedback on the suggestions and alerts to the server, contributing to improving the accuracy of the system.

[0368] Specific examples

[0369] The server takes into consideration the user's emotional state and notifies them of improvement suggestions at a time when they are least stressed.

[0370] The user checks the proposed new work procedure and instructs the on-site workers on the new procedure.

[0371] The user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[0372] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[0373] Prompt Sentence Examples

[0374] "Explain how you can collect data from 3D cameras and microphones installed in a factory, denoise the audio data, and analyze the emotional state."

[0375] "Please explain in detail the steps to train the AI ​​model using past accident data and work procedures."

[0376] "How can we collect real-time data from the field to optimize safety and efficiency?"

[0377] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0378] Step 1:

[0379] The server collects information from input devices installed on-site. Specifically, it receives data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. Inputs include video data, audio data, temperature data, and vibration data. The collected raw data is then output.

[0380] Step 2:

[0381] The server then performs a noise reduction process on the collected audio data. Specifically, it uses a speech recognition algorithm to filter out background noise and extract only the target voice. The input includes raw audio data, and the output is clear audio data with noise removed.

[0382] Step 3:

[0383] The server performs motion detection on the collected video data and removes frames in which no specific motion is detected. Specifically, it uses a video analysis algorithm to remove frames without motion. The input includes raw video data, and the output is video data with motion.

[0384] Step 4:

[0385] The server analyzes text data (such as accident reports, near-miss incident logs, and work procedures) using natural language processing (NLP) technology to extract important keywords and phrases. Specifically, it uses an NLP algorithm to extract keywords such as "fall," "fire," and "machine malfunction" from the text. The input includes the text data, and the output is the extracted keywords and phrases.

[0386] Step 5:

[0387] The server converts data of different formats into a unified format and generates a single integrated dataset. Specifically, it converts PDF-formatted work procedures into text format and integrates various sensor data along a time axis. Inputs include text data, audio data, video data, and sensor data, and the output is an integrated dataset.

[0388] Step 6:

[0389] The server uses the combined dataset to train the AI ​​model, specifically, using a deep learning algorithm to train the model, with the combined dataset as input and the trained AI model as output.

[0390] Step 7:

[0391] The server uses a trained AI model to analyze the collected real-time data. Specifically, it uses the AI ​​model to determine the efficiency and safety of the site. The input includes real-time data, and the output is an evaluation result regarding efficiency and safety.

[0392] Step 8:

[0393] The server generates improvement proposals and emergency alerts based on the evaluation results and notifies the user. Specifically, it uses AI to evaluate the analysis results and generate proposals and alerts as necessary. The input includes the evaluation results, and the generated improvement proposals and alerts are obtained as output.

[0394] Step 9:

[0395] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes audio and video data to identify the user's emotional state. The input includes audio and video data, and the output is an analysis result related to the user's emotional state.

[0396] Step 10:

[0397] The server provides appropriate feedback based on the user's emotional state. Specifically, it adjusts the content and timing of the feedback based on the user's emotional state. The input includes the analysis results of the user's emotional state, and the output is an appropriate feedback message.

[0398] (Application example 2)

[0399] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0400] Maintaining high levels of efficiency and safety at the same time is difficult in on-site work, requiring immediate responses to on-site conditions. Furthermore, lack of appropriate feedback that takes into account the emotional state of workers can lead to employee stress and anxiety that negatively impacts efficiency and safety. The purpose of this invention is to solve these problems and provide an excellent working environment.

[0401] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0402] In this invention, the server includes means for collecting information from input devices installed at the site, means for pre-processing the collected information, and means for integrating the pre-processed information to generate a data set, thereby optimizing the efficiency and safety of the site and providing appropriate feedback that takes into account the emotional state of the user.

[0403] "Input devices installed on-site" are devices installed in factories and construction sites to collect various information, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[0404] "Means for preprocessing collected information" refers to a function for cleansing and processing initial data, such as removing noise from collected audio data and analyzing video data.

[0405] The "means for generating a dataset" is a function that integrates pre-processed data of various formats and converts them into a single consistent format.

[0406] "Means of learning the model" refers to the ability to train an AI model based on an integrated dataset using techniques such as deep learning.

[0407] The "means for determining efficiency and safety" is a function that uses a trained AI model to analyze data collected in real time and evaluate the efficiency and safety of work.

[0408] The "means for generating and notifying improvement proposals and emergency alerts" is a function that generates optimal improvement proposals based on the judgment results and notifies the user of emergency alerts as necessary.

[0409] The "means including an emotion engine" is a function that analyzes the user's voice and video data to recognize the user's emotional state and provides feedback based on the emotional state.

[0410] "Past accident data" refers to data containing detailed information about accidents that have occurred in the past.

[0411] A "near miss incident diary" is a record of near miss incidents and close calls that have occurred in the past.

[0412] A "work procedure manual" is a document that describes the procedures and methods for performing a specific task.

[0413] "Performance data" refers to data relating to the amount of work and results achieved on-site within a specific period of time.

[0414] "Machine specification data" refers to data that includes detailed technical specifications of machines and equipment used on-site.

[0415] "Multimodal AI" refers to AI technology that has the ability to comprehensively analyze data in different formats (such as audio, video, text, etc.).

[0416] This invention relates to an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site and performs real-time analysis to identify inefficient work areas and dangerous spots, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes and analyzes the user's emotional state to provide appropriate feedback to the user based on their emotional state.

[0417] System Overview

[0418] This system has the following main functions:

[0419] 1. Information gathering

[0420] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, and other devices installed on-site.

[0421] The collected data includes past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0422] 2. Data Preprocessing

[0423] The server performs noise reduction processing on the collected audio data and deletes frames in which no specific movement is detected from the video data.

[0424] The text data is analyzed using natural language processing (NLP) techniques to extract important keywords and phrases.

[0425] Data of different formats is converted into a unified format.

[0426] 3. Data integration and model training

[0427] The server integrates the pre-processed data to create a single integrated dataset and stores it in a cloud database.

[0428] The integrated data set will be used to train AI models to improve efficiency and safety analysis capabilities.

[0429] 4. Introducing the Emotion Engine

[0430] To recognize the user's emotions, the server analyzes audio data obtained from a microphone and video data obtained from a camera to recognize the user's emotional state (anger, sadness, joy, etc.).

[0431] Providing appropriate feedback to the user based on the perceived emotional state.

[0432] 5. Real-time analysis and improvement suggestions

[0433] The server analyzes real-time data from the site and monitors work efficiency and safety.

[0434] Based on the analysis results, inefficient areas and potential danger areas are identified and appropriate improvement proposals are generated.

[0435] If an emergency occurs, an alert is sent to the user immediately.

[0436] 6. Notifications and Feedback

[0437] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the notifications based on the user's emotional state.

[0438] Users can check the suggestions and alerts, take necessary actions, and provide feedback on the suggestions and alerts to the server, thereby contributing to improving the accuracy of the system.

[0439] Hardware and software used

[0440] The system uses the following hardware and software:

[0441] Hardware: 3D camera, microphone, temperature sensor, vibration sensor, cloud database

[0442] Software: OpenCV, PyAudio, TensorFlow, HuggingFace Transformers, NLP tools

[0443] Specific examples

[0444] For example, if "fall accidents" occur frequently in a factory, the following prompts can be input into the generative AI model based on the collected data to obtain improvement suggestions:

[0445] Example prompt

[0446] Analyze all accident reports and sensor signal logs for tip-over accidents that have occurred over the past six months and generate specific proposals for improving safety. Improvement proposals may include real-time monitoring, changes to work procedures, and repositioning of machinery.

[0447] This makes the system a powerful tool for improving efficiency and safety on-site.

[0448] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0449] Step 1: Gather information

[0450] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. installed on-site. The input in this step is raw data from each sensor, and the output is an unprocessed data stream. Specifically, the server periodically polls data from each input device and transfers it to a cloud database.

[0451] Step 2: Data Preprocessing

[0452] The server performs noise reduction on the collected audio data and removes frames from the video data where no specific motion is detected. The input is the raw data collected in step 1, and the output is the pre-processed, clean data. Specifically, it applies a noise filter to the audio data and a frame removal algorithm to the video data.

[0453] Step 3: Data Integration

[0454] The server integrates preprocessed data in various formats to generate a single dataset. The input is preprocessed audio, video, temperature, vibration, and other data, and the output is a unified-format dataset. Specifically, the server links data in different formats using timestamps and stores them in a cloud database as a single integrated dataset.

[0455] Step 4: Model training

[0456] The server uses the dataset to train the AI ​​model. The input is the dataset generated in step 3, and the output is the trained AI model. Specifically, the server uses a deep learning algorithm to train the data. Here, libraries such as TensorFlow are used to train the model.

[0457] Step 5: Real-time analysis

[0458] The server analyzes real-time data from the site to determine efficiency and safety. The input is the data collected in real time and the trained model, and the output is the evaluation results of efficiency and safety. Specifically, the server uses the AI ​​model to analyze the real-time data and identify inefficient areas and potential dangerous areas.

[0459] Step 6: Sentiment Analysis

[0460] The server uses an emotion engine to analyze the user's emotions and recognize their emotional state from audio and video data. The input is the user's audio and video data, and the output is the evaluation result of their emotional state. Specifically, the emotion engine uses voice recognition and facial expression recognition algorithms to analyze the user's emotional state.

[0461] Step 7: Generate improvement suggestions and emergency alerts

[0462] The server generates and notifies improvement proposals and emergency alerts based on the analysis results. The input is the analysis results from steps 5 and 6, and the output is improvement proposals and emergency alerts. Specifically, the server generates optimal improvement proposals based on the analysis results and notifies the user device, such as a smartphone or tablet.

[0463] Step 8: Gather feedback and fine-tune the system

[0464] The user reviews the suggestions and alerts and provides their feedback to the server. The input is the user's feedback, and the output is the system's adjustment results. Specifically, the user enters their evaluation of the suggestions and alerts into a feedback form, and the server uses this information to improve the accuracy of the AI ​​model and analysis algorithms.

[0465] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0466] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0467] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0468] [Second embodiment]

[0469] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0470] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0471] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0472] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0473] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0474] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0475] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0476] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0477] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0478] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0479] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0480] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0481] This invention describes an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement proposals and emergency alerts. The operation of this system's program is explained below in natural language.

[0482] Data Collection Overview

[0483] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[0484] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[0485] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0486] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[0487] Data Preprocessing Overview

[0488] 3. The server performs preprocessing on the collected data, specifically removing noise from the audio data, deleting unnecessary frames from the video data, and analyzing the text data.

[0489] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[0490] 4. The server converts the different types of data into a unified format.

[0491] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[0492] Overview of Data Integration and Model Training

[0493] 5. The server aggregates the pre-processed data to create a single aggregated dataset.

[0494] Example: The server stores cleaned video data, audio data, sensor data, and text data as a single dataset.

[0495] 6. The server trains the multimodal AI model using the integrated dataset.

[0496] Example: The server uses a deep learning framework to train models and improve its efficiency and safety decision-making capabilities.

[0497] Real-time analysis and improvement suggestions overview

[0498] 7. The server analyzes the collected real-time data and monitors the situation on site.

[0499] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0500] 8. Based on the analysis results, the server identifies inefficient areas and potential dangers and generates appropriate improvement proposals.

[0501] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[0502] 9. The server will immediately alert the user if an imminent danger occurs.

[0503] Example: The server detects signs of a fire from video data and immediately sends a warning message to the administrator's terminal.

[0504] Notifications and Feedback Overview

[0505] 10. The server notifies the user of the generated improvement suggestions and emergency alerts.

[0506] Example: The server sends a message to the user's device saying, "Introducing a new work procedure will improve efficiency."

[0507] 11. The user checks the notified suggestions and alerts and takes the necessary action.

[0508] Example: A user reviews proposed work procedure changes and instructs field workers on the new procedures.

[0509] 12. Users provide feedback to the server on suggestions and alerts, helping to improve the system's accuracy.

[0510] Example: A user fills in a feedback form with a newly encountered problem and submits it to the server.

[0511] The above is a specific embodiment of the system of the present invention, which can significantly improve efficiency and safety in factories and construction sites.

[0512] The processing flow will be explained below.

[0513] Step 1:

[0514] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[0515] Step 2:

[0516] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[0517] Step 3:

[0518] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements. For example, it deletes frames in which human movement cannot be confirmed.

[0519] Step 4:

[0520] The server analyzes text data (accident reports, near-miss incident logs, work procedure manuals, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[0521] Step 5:

[0522] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0523] Step 6:

[0524] The server then combines the pre-processed data to create a single integrated dataset, for example audio, video, temperature, and vibration data, which is stored in a cloud database and prepared for analysis.

[0525] Step 7:

[0526] The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0527] Step 8:

[0528] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal conditions from real-time video and extracting abnormal sounds from audio data.

[0529] Step 9:

[0530] The server uses the analysis results to identify areas of inefficiency and potential dangers. For example, if a location where workers frequently stop or excessive vibration is detected, it will identify that as a potential danger.

[0531] Step 10:

[0532] The server generates improvement suggestions aimed at improving efficiency and safety, for example by changing the layout of machines or proposing new work procedures.

[0533] Step 11:

[0534] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[0535] Step 12:

[0536] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[0537] Step 13:

[0538] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0539] Example 1

[0540] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0541] In factories and construction sites, real-time data collection and analysis is necessary to ensure both work efficiency and safety, but existing systems are difficult to adequately address this. There is also a need to integrate data from a variety of sensors and past data, and quickly generate appropriate improvement proposals and emergency alerts. Conventional technologies face challenges in unifying different data formats, removing noise, and improving the accuracy of real-time analysis.

[0542] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0543] In this invention, the server includes means for collecting data from input devices installed on-site, preprocessing means for removing noise and unnecessary data from the collected data, means for converting data of different formats into a unified format, means for integrating the preprocessed data to generate a dataset, means for training a multimodal model using the dataset, means for analyzing real-time data to determine efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting feedback from users to improve the accuracy of the system, thereby enabling significant improvements in work efficiency and safety at factories and construction sites.

[0544] A "server" is a central processing system that collects, processes, and analyzes data from various input devices installed in factories and construction sites, and notifies and presents the results to users.

[0545] "Input devices" are devices installed on-site to collect data, and include 3D cameras, microphones, temperature sensors, vibration sensors, etc.

[0546] The "data collection means" is a mechanism for continuously acquiring data from the input device and transmitting it to the server.

[0547] "Noise reduction means" refers to the algorithms and filtering processes used to remove background noise from collected audio data.

[0548] "Means for removing unnecessary data" refers to steps for eliminating frames in which no specific movement is detected or meaningless information from the video data.

[0549] "Preprocessing means" refers to a technique that includes processes such as noise removal and deletion of unnecessary data on collected raw data to make it ready for analysis.

[0550] "Means of converting into a unified format" refers to technology that unifies data collected in different formats into a consistent format, enabling integrated analysis.

[0551] "Dataset generation means" refers to the process of integrating preprocessed data and generating and saving a single integrated dataset.

[0552] A "multimodal model" is an artificial intelligence model that simultaneously analyzes different types of data (e.g., video, audio, text, sensor information) and learns their interrelationships.

[0553] "Model training methods" are techniques for training AI models using preprocessed and integrated datasets to improve their accuracy.

[0554] "Real-time data analysis means" refers to technology that instantly analyzes collected data and evaluates and monitors the current on-site situation.

[0555] "Means for determining efficiency and safety" refers to an algorithm that evaluates on-site work efficiency and safety based on the results of analyzing real-time data and makes specific judgments.

[0556] The "improvement proposal generation means" is a process for generating specific proposals for improving work efficiency and safety based on the results of data analysis.

[0557] The "emergency alert notification means" is a mechanism for immediately issuing a warning to the user when a serious danger is detected.

[0558] The "feedback collection means" is a mechanism for collecting user opinions and suggestions for improvement on a server and using them to improve and enhance the accuracy of the system.

[0559] The present invention relates to a system for optimizing the efficiency and safety of work in factories and construction sites. This system is composed of a server, an input device, and a user terminal.

[0560] The server collects data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. This data includes video data, audio data, temperature data, and vibration data. For example, the 3D camera streams video of the site in real time, and the microphone captures audio from the site. The temperature and vibration sensors collect their respective environmental data.

[0561] The server removes noise from the collected data and deletes unnecessary data. This preprocessing improves the accuracy of the analysis. Specifically, a voice recognition algorithm is used to remove noise from the audio data, and a motion detection algorithm is used to delete unnecessary frames from the video data. In addition, PDF-formatted work procedures and past accident reports are converted into text format using OCR technology.

[0562] The preprocessed data is then converted into a unified format by the server. This conversion integrates data from different formats into a consistent format. The converted data is then stored on the server as a unified dataset and analyzed using a multimodal model.

[0563] A multimodal model is an artificial intelligence model that simultaneously analyzes different types of data, optimally analyzing different modalities (video, audio, text, and sensor information). The server uses this multimodal model to determine efficiency and safety. For example, real-time video analysis can detect worker movements and identify unnatural movements. Audio analysis can also detect abnormal sounds on-site.

[0564] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and proposes relocation of the machines. In addition, if an emergency alert is required, the server immediately sends a warning message to the user. For example, it detects signs of a fire from video data and immediately sends a warning message to the user's device.

[0565] Users check improvement suggestions and emergency alerts notified by the server. Based on the notified information, they can change work procedures or instruct on-site workers on emergency responses. Users can also provide feedback on suggestions and alerts to the server, contributing to improving the accuracy of the system. For example, users can enter newly discovered problems or areas for improvement into a feedback form and send it to the server.

[0566] As a result, the efficiency and safety of work at factories and construction sites can be significantly improved. The system of the present invention can quickly and accurately monitor the situation at the site by analyzing real-time data, and provide appropriate improvement suggestions and emergency alerts.

[0567] An example of a prompt is, "Please analyze the current factory data and generate suggestions to improve safety." Using this prompt, the generative AI model can make appropriate analyses and suggestions.

[0568] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0569] Step 1:

[0570] The server collects data from input devices installed on-site. These include a 3D camera, microphone, temperature sensor, and vibration sensor. The server receives a real-time video stream from the 3D camera and simultaneously captures audio data from the microphone. Data from the temperature sensor and vibration sensor is also sent to the server. This allows for an immediate understanding of the situation on-site.

[0571] Input: 3D camera image data, microphone audio data, temperature sensor and vibration sensor data

[0572] Output: Collected raw data (video, audio, temperature, vibration)

[0573] Step 2:

[0574] The server performs preprocessing on the collected data. Specifically, it removes noise from the audio data and deletes unnecessary frames from the video data. To remove noise from the audio data, it uses a speech recognition algorithm to remove background noise. For the video data, it uses a motion detection algorithm to remove inactive frames. Furthermore, it converts work procedures and accident reports into text data using OCR.

[0575] Input: Collected raw data (video, audio, temperature, vibration), work procedures, accident reports

[0576] Output: Preprocessed data (noise-removed audio, video with unnecessary frames deleted, materials converted into text data)

[0577] Step 3:

[0578] The server converts the pre-processed data into a unified format, which aligns data from different formats into a consistent format, such as text data converted from PDF or sensor data converted to JSON format.

[0579] Input: Preprocessed data

[0580] Output: Dataset converted to a unified format

[0581] Step 4:

[0582] The server then combines the data in a unified format to generate a single comprehensive dataset. All data is integrated based on timestamps and stored in a database. This dataset is then used for model training and real-time analysis.

[0583] Input: Data converted into a unified format

[0584] Output: Unified dataset

[0585] Step 5:

[0586] The server trains a multimodal AI model using the integrated dataset. It uses a deep learning framework (e.g., TensorFlow) to analyze patterns in the data and improve its ability to make safety and efficiency decisions. This training allows the model to analyze different types of data simultaneously.

[0587] Input: Integrated dataset

[0588] Output: Trained multimodal AI model

[0589] Step 6:

[0590] The server analyzes the collected real-time data and monitors the situation on-site. Real-time video analysis detects worker movements and identifies unnatural movements. Audio analysis detects abnormal sounds.

[0591] Input: Real-time data (video, audio, temperature, vibration)

[0592] Output: Analysis results (on-site condition monitoring data)

[0593] Step 7:

[0594] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and suggests changing the machine's location. It also immediately sends a warning message to the user if an emergency alert is required. For example, it detects signs of a fire from video data and sends a warning message to the user's device.

[0595] Input: Analysis results

[0596] Output: Improvement suggestions and emergency alerts

[0597] Step 8:

[0598] The server notifies the generated improvement suggestions and emergency alerts to the user's device, which has the function to receive and display the suggestions and alerts.

[0599] Input: Improvement suggestions and emergency alerts

[0600] Output: Notification to user terminal

[0601] Step 9:

[0602] The user checks the notified improvement proposals and alerts and takes the necessary action. For example, the user checks the proposed changes to work procedures and instructs the on-site workers on the new procedures. In the case of an emergency alert, the user can quickly implement countermeasures.

[0603] Input: Improvement suggestions and emergency alerts

[0604] Output: User response

[0605] Step 10:

[0606] Users can provide feedback on suggestions and alerts to the server, helping to improve the system's accuracy. This feedback is reflected in the next model training.

[0607] Input: User feedback

[0608] Output: Feedback data to the server

[0609] The above is the specific processing flow of this system.

[0610] (Application example 1)

[0611] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0612] To maximize the efficiency and safety of on-site work, it is important to monitor the situation in real time and quickly and accurately identify inefficient or dangerous areas. However, conventional systems have difficulty integrating dispersed data and analyzing it in real time, resulting in low accuracy in improvement proposals and emergency alerts. Furthermore, a monitoring system using advanced data analysis technology was required to enable robots to work safely and effectively on-site.

[0613] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0614] In this invention, the server includes means for collecting information from input devices installed on-site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for training a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting, analyzing, and monitoring robot sensor data in real time. This makes it possible to optimize the efficiency and safety of on-site work, reducing risks and improving work safety, particularly at sites where robots work.

[0615] "Worksite" refers to a place where work is carried out, such as a factory or construction site.

[0616] "Input devices" refers to devices such as sensors and cameras installed on-site to collect data.

[0617] "Preprocessing" refers to the process of removing unnecessary parts from collected data and converting it into a format that is easy to analyze.

[0618] "Dataset" refers to a set of data that is collected, preprocessed, and integrated from multiple data sources.

[0619] A "model" is an algorithm that learns from collected and preprocessed data to accomplish a specific task.

[0620] "Real-time data" refers to data that is collected continuously along an ongoing timeline.

[0621] "Efficiency" refers to the degree to which work is carried out smoothly and without waste.

[0622] "Safety" refers to the degree to which work is carried out without hazards.

[0623] "Improvement proposals" refer to recommendations for improving efficiency or safety.

[0624] "Emergency Alert" means an emergency notification of immediate danger at a site.

[0625] "Robot" refers to a mechanical device that is programmed to perform tasks automatically.

[0626] "Sensor data" refers to data such as temperature, vibration, audio, and video collected by sensors.

[0627] "Monitoring" refers to the act of continuously observing the situation at a site and detecting any changes.

[0628] "Camera" refers to a device that captures images and collects them as digital data.

[0629] "Microphone" refers to a device for collecting sound.

[0630] "Temperature sensor" refers to a device that measures temperature and collects it as digital data.

[0631] A "vibration sensor" refers to a device that detects vibrations and collects them as digital data.

[0632] "Real-time" means happening at the present time, without delay.

[0633] "Notification" refers to the act of sending a message to convey information.

[0634] "Server" refers to a computer system for collecting, preprocessing, integrating, analyzing, and notifying data.

[0635] This invention is realized using an AI system to optimize the efficiency and safety of work on-site. The system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient areas and dangerous spots, and generates improvement suggestions and emergency alerts.

[0636] 1. Program Overview

[0637] The server collects data from input devices such as 3D cameras, microphones, temperature sensors, and vibration sensors. It then preprocesses and integrates the collected data to generate a dataset. This dataset is used to train an AI model that analyzes real-time on-site data. Based on the analysis results, it makes efficiency and safety judgments, generates improvement suggestions and emergency alerts, and notifies the user's device or the robot's display.

[0638] 2. Hardware and Software Used

[0639] 3D camera: Collects footage of the scene in real time.

[0640] Microphone: Collects audio data.

[0641] Temperature Sensor: Collects temperature data in the field.

[0642] Vibration Sensor: Collects vibration data.

[0643] Server: Collects, preprocesses, integrates, analyzes, trains models, and notifies data. In particular, it uses deep learning models using TensorFlow and Keras.

[0644] User device: Receives and displays improvement suggestions and emergency alerts.

[0645] Robot: Monitors on-site operations and presents the necessary information to users in an easy-to-understand format.

[0646] 3. Data processing and calculation

[0647] The server performs the following data processing.

[0648] Data collection: Collect data in real time from 3D cameras, microphones, temperature sensors, and vibration sensors.

[0649] Data preprocessing: Converting video data to grayscale, removing noise from audio data, detecting outliers in temperature and vibration data, etc.

[0650] Data integration: Integrate various pre-processed data to generate a single dataset.

[0651] Model training: Use TensorFlow and Keras to train a multimodal AI model based on the integrated dataset.

[0652] Real-time analysis: Trained models analyze collected real-time data to assess efficiency and safety.

[0653] Notification generation: Based on the analysis results, improvement suggestions and emergency alerts are automatically generated and sent to the user's device.

[0654] 4. Examples and prompts

[0655] A concrete example of this system in action is a robot in a factory. The robot operates safely by collecting data from 3D cameras and temperature sensors in the work area. The server analyzes this data in real time, identifies abnormal temperature rises or unnatural movements in the work area, and immediately sends an emergency alert to the manager's terminal.

[0656] Example prompt sentence:

[0657] "Please explain in detail how real-time data from 3D cameras and temperature sensors is analyzed to train AI models that determine efficiency and risk so that factory robots can continue to work safely."

[0658] By implementing this invention, the efficiency and safety of work in factories and construction sites can be significantly improved.

[0659] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0660] Step 1:

[0661] The server collects data from 3D cameras, microphones, temperature sensors, and vibration sensors installed on-site. Video data, audio data, temperature data, and vibration data obtained from input devices are input to the server. The collected data is in the form of a video stream for video, an audio stream for audio, and numerical data output from the sensors for temperature and vibration. This allows the real-time situation on-site to be accumulated on the server as digital data.

[0662] Step 2:

[0663] The server performs preprocessing on the collected data. This involves converting the video data to grayscale and removing noise from the audio data. It also detects and filters out abnormal values ​​in the temperature and vibration data. Specifically, OpenCV is used to convert each frame of the video data to grayscale, and an audio filtering algorithm is used to remove unwanted noise from the audio data. This results in clean data suitable for analysis.

[0664] Step 3:

[0665] The server then combines the pre-processed data into a single dataset, including video frames converted to grayscale, filtered audio data, and normalized temperature and vibration data, consolidating the various data types into a single, integrated format for further analysis.

[0666] Step 4:

[0667] The server trains an AI model using the integrated dataset. Specifically, it uses TensorFlow and Keras to train a deep learning model. The input data is a preprocessed and integrated dataset, and the output is a model for evaluating site efficiency and safety. This model learns based on historical data and real-time input data.

[0668] Step 5:

[0669] The server uses a trained AI model to analyze on-site data collected in real time. The input data is video data, audio data, temperature data, and vibration data collected in real time, and the model analyzes this data to evaluate efficiency and safety. The analysis results output a judgment that a particular task is inefficient or dangerous.

[0670] Step 6:

[0671] The server generates improvement suggestions and emergency alerts based on the analysis results. For example, if workers' movements in a specific area are unnatural, it will suggest changes to work procedures as an improvement suggestion. If temperature data is abnormally high, it will determine that there is a risk of fire and generate an emergency alert. These notifications are generated in a specific and actionable format.

[0672] Step 7:

[0673] The server notifies the generated improvement proposals and emergency alerts to the user's terminal or the robot's display. The notification content is displayed so that the user can immediately understand it and take measures. Specifically, the server sends email notifications using SMTP and displays warning messages on the robot's display.

[0674] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0675] This invention describes an AI system for optimizing the efficiency and safety of work in factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. Furthermore, this invention combines an emotion engine that recognizes the user's emotions, enabling it to provide appropriate feedback based on the user's emotional state.

[0676] Data Collection Overview

[0677] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. If there is a shortage, it also refers to backup data.

[0678] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[0679] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0680] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[0681] Data Preprocessing Overview

[0682] 3. The server performs noise reduction on the collected audio data, specifically filtering out background noise and extracting only the target audio.

[0683] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[0684] 4. The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[0685] 5. The server converts different data formats into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0686] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[0687] Overview of Data Integration and Model Training

[0688] 6. The server combines the pre-processed data to create a single combined data set, for example audio, video, temperature, and vibration data, and stores it in a cloud database, ready for analysis.

[0689] 7. The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0690] Example: The server uses past data to train an AI model to help determine future work efficiency and safety.

[0691] Introducing the Emotion Engine

[0692] 8. The server combines an emotion engine that recognizes the user's emotions and analyzes the user's emotional state. For example, it analyzes audio data acquired from a microphone and recognizes the user's emotional state (anger, sadness, joy, etc.).

[0693] Example: The server uses speech recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[0694] 9. The server analyzes the user's facial expressions from the video data and detects changes in emotions. For example, it analyzes video data acquired from a camera and determines whether the user is smiling or confused.

[0695] Example: The server uses a facial expression recognition algorithm to suggest appropriate actions to take if the user is feeling anxious.

[0696] Real-time analysis and improvement suggestions overview

[0697] 10. The server analyzes the collected real-time data and monitors the situation at the site. It detects worker movements and abnormal situations from real-time video and extracts abnormal sounds from audio data.

[0698] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0699] 11. Based on the analysis results, the server identifies inefficiencies and potential dangers and generates appropriate improvement proposals, such as changing the layout of machines or proposing new work procedures.

[0700] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[0701] 12. The server will immediately send an alert to the user if an emergency occurs. If it detects signs of a fire, it will immediately send a warning message to the administrator's terminal.

[0702] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[0703] Notifications and Feedback Overview

[0704] 13. The server notifies the user of the generated improvement suggestions and emergency alerts. The server adjusts the content and timing of the feedback based on the user's emotional state.

[0705] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[0706] 14. The user checks the suggestions and alerts sent from the server and takes necessary action. For example, the user communicates the proposed new work procedures to the field workers and instructs them to carry them out immediately.

[0707] Example: A user reviews the proposed new work procedure and instructs field workers on the new procedure.

[0708] 15. Users contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0709] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[0710] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[0711] The processing flow will be explained below.

[0712] Step 1:

[0713] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[0714] Step 2:

[0715] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[0716] Example: A server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone.

[0717] Step 3:

[0718] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements.

[0719] Example: The server removes still frames from video data and extracts only frames with movement.

[0720] Step 4:

[0721] The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases.

[0722] Example: The server analyzes accident reports and extracts important keywords such as "fall," "fire," and "machine malfunction."

[0723] Step 5:

[0724] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0725] Example: A server converts PDF files into text format to create a dataset for multimodal AI.

[0726] Step 6:

[0727] The server aggregates the pre-processed data to create a single aggregated data set.

[0728] Example: The server combines cleaned audio, video, sensor, and text data to create a unified dataset.

[0729] Step 7:

[0730] The server uses the combined data set to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0731] Example: The server uses past data to train AI models to improve future work efficiency and safety.

[0732] Step 8:

[0733] The server uses an emotion engine that recognizes the user's emotions to analyze the user's voice data and recognize the user's emotional state (anger, sadness, joy, etc.).

[0734] Example: A server uses voice recognition and emotion analysis algorithms to determine whether a user is stressed.

[0735] Step 9:

[0736] The server analyzes the user's facial expressions from the video data and detects changes in emotions.

[0737] Example: The server uses a facial expression recognition algorithm to determine whether the user is smiling, confused, etc.

[0738] Step 10:

[0739] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal situations from real-time video and extracting abnormal sounds from audio data.

[0740] Example: The server uses real-time video and audio analysis to detect worker movements and identify unnatural movements.

[0741] Step 11:

[0742] Based on the analysis results, the server identifies inefficient areas and potential danger points and generates appropriate improvement proposals.

[0743] Example: The server analyzes the locations where workers frequently stop and suggests relocating those locations.

[0744] Step 12:

[0745] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[0746] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[0747] Step 13:

[0748] The server notifies the user of the generated improvement suggestions and emergency alerts, and adjusts the content and timing of the feedback based on the user's emotional state.

[0749] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[0750] Step 14:

[0751] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[0752] Example: A user reviews a proposed new work procedure and communicates the new procedure to workers in the field.

[0753] Step 15:

[0754] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0755] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[0756] Example 2

[0757] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0758] Optimizing the efficiency and safety of on-site work requires real-time situational awareness and appropriate feedback. However, existing systems lack the ability to integrate and analyze collected data, making it difficult to respond quickly and accurately to specific issues. In addition, they are unable to provide feedback that takes into account the emotional state of workers, limiting further improvements in efficiency and safety.

[0759] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting information from input devices installed at the site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for learning a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for analyzing the emotional state of the user and providing appropriate feedback based on the emotional state. This makes it possible to analyze the situation at the site in real time, improve efficiency and safety, and provide appropriate feedback based on the emotional state of the worker.

[0760] "Worksite" refers to the location where work actually takes place, such as a factory or construction site.

[0761] "Input devices" refers to equipment including sensors, cameras, microphones, etc. that are installed on-site to collect information.

[0762] "Means for collecting information" refers to the method or technology for transmitting data obtained from the input device to the server.

[0763] "Preprocessing" refers to a series of steps taken to convert collected raw data into a form that is easier to analyze.

[0764] "Preprocessing means" refers to techniques and methods for performing preprocessing, such as noise removal, data transformation, and extraction of important information.

[0765] A "dataset" refers to a single set of data that is created by integrating multiple preprocessed data.

[0766] "Means for generating datasets" refers to methods and techniques for combining pre-processed data into a format that an AI model can learn from.

[0767] "Means of training the model" refers to the methods and techniques used to train the AI ​​model using the generated dataset.

[0768] "Real-time data" refers to the latest data collected from the field.

[0769] "Means of analysis" refers to the techniques and methods used to analyze collected data and determine its efficiency and safety.

[0770] "Decision result" refers to the conclusion or evaluation obtained based on the analyzed data.

[0771] "Improvement proposals" refer to specific proposals or methods for improving efficiency or safety.

[0772] An "urgent alert" is a warning issued when a danger or problem occurs that requires immediate action.

[0773] "Means for notification" refers to the methods and technologies for communicating generated improvement suggestions and emergency alerts to users.

[0774] "Emotional state" refers to a user's mental state or emotion, including, for example, joy, anger, sadness, and the like.

[0775] "Means for providing feedback" refers to methods and techniques for providing appropriate information or advice to a user based on the analyzed emotional state.

[0776] This invention is an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes the user's emotional state to provide appropriate feedback.

[0777] The system configuration is as follows:

[0778] Data collection

[0779] The server collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. For example, the server receives video streams from 3D cameras and simultaneously captures audio data from microphones. It also collects data on past accidents, near-miss incident logs, work procedures, completed volume data, and machine specification data.

[0780] Specific examples

[0781] The server simultaneously collects the video stream from the 3D camera and the audio data from the microphone.

[0782] The server retrieves past fall accident reports and work procedure manuals in PDF format and stores them in a database.

[0783] Data Preprocessing

[0784] The server performs noise reduction on the collected audio data. Specifically, it filters out background noise from the audio data and extracts only the target audio. It also removes frames from the video data in which no specific movement is detected.

[0785] In addition, the server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, and also converts data in different formats into a unified format.

[0786] Specific examples

[0787] The server uses a voice recognition algorithm to remove background noise from the voice data acquired from the microphone.

[0788] The server converts the PDF files into text format, generating a dataset that is easy to analyze.

[0789] Data integration and model training

[0790] The server combines the pre-processed data to create a single integrated data set, which is then used to train an AI model that uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0791] Specific examples

[0792] The server uses past data to train the AI ​​model, helping to determine future work efficiency and safety.

[0793] Introducing the Emotion Engine

[0794] The server combines an emotion engine to recognize the user's emotional state. For example, it analyzes audio data acquired from a microphone to recognize the user's emotional state (anger, sadness, joy, etc.). It also analyzes the user's facial expressions from video data to detect emotional fluctuations.

[0795] Specific examples

[0796] The server uses voice recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[0797] The server uses a facial expression recognition algorithm to suggest appropriate measures to address the user's anxiety.

[0798] Real-time analysis and improvement suggestions

[0799] The server analyzes the collected real-time data and monitors the situation at the site. Based on the analysis results, it identifies inefficiencies and potential dangers, generates improvement proposals, and immediately sends alerts if an emergency danger occurs, if necessary.

[0800] Specific examples

[0801] The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0802] If the server detects a fire, it will send a real-time alert to the administrator's smartphone.

[0803] Notifications and Feedback

[0804] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the feedback based on the user's emotional state.

[0805] The user checks the suggestions and alerts sent from the server and takes necessary action. The user also provides feedback on the suggestions and alerts to the server, contributing to improving the accuracy of the system.

[0806] Specific examples

[0807] The server takes into consideration the user's emotional state and notifies them of improvement suggestions at a time when they are least stressed.

[0808] The user checks the proposed new work procedure and instructs the on-site workers on the new procedure.

[0809] The user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[0810] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[0811] Prompt Sentence Examples

[0812] "Explain how you can collect data from 3D cameras and microphones installed in a factory, denoise the audio data, and analyze the emotional state."

[0813] "Please explain in detail the steps to train the AI ​​model using past accident data and work procedures."

[0814] "How can we collect real-time data from the field to optimize safety and efficiency?"

[0815] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0816] Step 1:

[0817] The server collects information from input devices installed on-site. Specifically, it receives data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. Inputs include video data, audio data, temperature data, and vibration data. The collected raw data is then output.

[0818] Step 2:

[0819] The server then performs a noise reduction process on the collected audio data. Specifically, it uses a speech recognition algorithm to filter out background noise and extract only the target voice. The input includes raw audio data, and the output is clear audio data with noise removed.

[0820] Step 3:

[0821] The server performs motion detection on the collected video data and removes frames in which no specific motion is detected. Specifically, it uses a video analysis algorithm to remove frames without motion. The input includes raw video data, and the output is video data with motion.

[0822] Step 4:

[0823] The server analyzes text data (such as accident reports, near-miss incident logs, and work procedures) using natural language processing (NLP) technology to extract important keywords and phrases. Specifically, it uses an NLP algorithm to extract keywords such as "fall," "fire," and "machine malfunction" from the text. The input includes the text data, and the output is the extracted keywords and phrases.

[0824] Step 5:

[0825] The server converts data of different formats into a unified format and generates a single integrated dataset. Specifically, it converts PDF-formatted work procedures into text format and integrates various sensor data along a time axis. Inputs include text data, audio data, video data, and sensor data, and the output is an integrated dataset.

[0826] Step 6:

[0827] The server uses the combined dataset to train the AI ​​model, specifically, using a deep learning algorithm to train the model, with the combined dataset as input and the trained AI model as output.

[0828] Step 7:

[0829] The server uses a trained AI model to analyze the collected real-time data. Specifically, it uses the AI ​​model to determine the efficiency and safety of the site. The input includes real-time data, and the output is an evaluation result regarding efficiency and safety.

[0830] Step 8:

[0831] The server generates improvement proposals and emergency alerts based on the evaluation results and notifies the user. Specifically, it uses AI to evaluate the analysis results and generate proposals and alerts as necessary. The input includes the evaluation results, and the generated improvement proposals and alerts are obtained as output.

[0832] Step 9:

[0833] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes audio and video data to identify the user's emotional state. The input includes audio and video data, and the output is an analysis result related to the user's emotional state.

[0834] Step 10:

[0835] The server provides appropriate feedback based on the user's emotional state. Specifically, it adjusts the content and timing of the feedback based on the user's emotional state. The input includes the analysis results of the user's emotional state, and the output is an appropriate feedback message.

[0836] (Application example 2)

[0837] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0838] Maintaining high levels of efficiency and safety at the same time is difficult in on-site work, requiring immediate responses to on-site conditions. Furthermore, lack of appropriate feedback that takes into account the emotional state of workers can lead to employee stress and anxiety that negatively impacts efficiency and safety. The purpose of this invention is to solve these problems and provide an excellent working environment.

[0839] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0840] In this invention, the server includes means for collecting information from input devices installed at the site, means for pre-processing the collected information, and means for integrating the pre-processed information to generate a data set, thereby optimizing the efficiency and safety of the site and providing appropriate feedback that takes into account the emotional state of the user.

[0841] "Input devices installed on-site" are devices installed in factories and construction sites to collect various information, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[0842] "Means for preprocessing collected information" refers to a function for cleansing and processing initial data, such as removing noise from collected audio data and analyzing video data.

[0843] The "means for generating a dataset" is a function that integrates pre-processed data of various formats and converts them into a single consistent format.

[0844] "Means of learning the model" refers to the ability to train an AI model based on an integrated dataset using techniques such as deep learning.

[0845] The "means for determining efficiency and safety" is a function that uses a trained AI model to analyze data collected in real time and evaluate the efficiency and safety of work.

[0846] The "means for generating and notifying improvement proposals and emergency alerts" is a function that generates optimal improvement proposals based on the judgment results and notifies the user of emergency alerts as necessary.

[0847] The "means including an emotion engine" is a function that analyzes the user's voice and video data to recognize the user's emotional state and provides feedback based on the emotional state.

[0848] "Past accident data" refers to data containing detailed information about accidents that have occurred in the past.

[0849] A "near miss incident diary" is a record of near miss incidents and close calls that have occurred in the past.

[0850] A "work procedure manual" is a document that describes the procedures and methods for performing a specific task.

[0851] "Performance data" refers to data relating to the amount of work and results achieved on-site within a specific period of time.

[0852] "Machine specification data" refers to data that includes detailed technical specifications of machines and equipment used on-site.

[0853] "Multimodal AI" refers to AI technology that has the ability to comprehensively analyze data in different formats (such as audio, video, text, etc.).

[0854] This invention relates to an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site and performs real-time analysis to identify inefficient work areas and dangerous spots, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes and analyzes the user's emotional state to provide appropriate feedback to the user based on their emotional state.

[0855] System Overview

[0856] This system has the following main functions:

[0857] 1. Information gathering

[0858] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, and other devices installed on-site.

[0859] The collected data includes past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0860] 2. Data Preprocessing

[0861] The server performs noise reduction processing on the collected audio data and deletes frames in which no specific movement is detected from the video data.

[0862] The text data is analyzed using natural language processing (NLP) techniques to extract important keywords and phrases.

[0863] Data of different formats is converted into a unified format.

[0864] 3. Data integration and model training

[0865] The server integrates the pre-processed data to create a single integrated dataset and stores it in a cloud database.

[0866] The integrated data set will be used to train AI models to improve efficiency and safety analysis capabilities.

[0867] 4. Introducing the Emotion Engine

[0868] To recognize the user's emotions, the server analyzes audio data obtained from a microphone and video data obtained from a camera to recognize the user's emotional state (anger, sadness, joy, etc.).

[0869] Providing appropriate feedback to the user based on the perceived emotional state.

[0870] 5. Real-time analysis and improvement suggestions

[0871] The server analyzes real-time data from the site and monitors work efficiency and safety.

[0872] Based on the analysis results, inefficient areas and potential danger areas are identified and appropriate improvement proposals are generated.

[0873] If an emergency occurs, an alert is sent to the user immediately.

[0874] 6. Notifications and Feedback

[0875] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the notifications based on the user's emotional state.

[0876] Users can check the suggestions and alerts, take necessary actions, and provide feedback on the suggestions and alerts to the server, thereby contributing to improving the accuracy of the system.

[0877] Hardware and software used

[0878] The system uses the following hardware and software:

[0879] Hardware: 3D camera, microphone, temperature sensor, vibration sensor, cloud database

[0880] Software: OpenCV, PyAudio, TensorFlow, HuggingFace Transformers, NLP tools

[0881] Specific examples

[0882] For example, if "fall accidents" occur frequently in a factory, the following prompts can be input into the generative AI model based on the collected data to obtain improvement suggestions:

[0883] Example prompt

[0884] Analyze all accident reports and sensor signal logs for tip-over accidents that have occurred over the past six months and generate specific proposals for improving safety. Improvement proposals may include real-time monitoring, changes to work procedures, and repositioning of machinery.

[0885] This makes the system a powerful tool for improving efficiency and safety on-site.

[0886] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0887] Step 1: Gather information

[0888] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. installed on-site. The input in this step is raw data from each sensor, and the output is an unprocessed data stream. Specifically, the server periodically polls data from each input device and transfers it to a cloud database.

[0889] Step 2: Data Preprocessing

[0890] The server performs noise reduction on the collected audio data and removes frames from the video data where no specific motion is detected. The input is the raw data collected in step 1, and the output is the pre-processed, clean data. Specifically, it applies a noise filter to the audio data and a frame removal algorithm to the video data.

[0891] Step 3: Data Integration

[0892] The server integrates preprocessed data in various formats to generate a single dataset. The input is preprocessed audio, video, temperature, vibration, and other data, and the output is a unified-format dataset. Specifically, the server links data in different formats using timestamps and stores them in a cloud database as a single integrated dataset.

[0893] Step 4: Model training

[0894] The server uses the dataset to train the AI ​​model. The input is the dataset generated in step 3, and the output is the trained AI model. Specifically, the server uses a deep learning algorithm to train the data. Here, libraries such as TensorFlow are used to train the model.

[0895] Step 5: Real-time analysis

[0896] The server analyzes real-time data from the site to determine efficiency and safety. The input is the data collected in real time and the trained model, and the output is the evaluation results of efficiency and safety. Specifically, the server uses the AI ​​model to analyze the real-time data and identify inefficient areas and potential dangerous areas.

[0897] Step 6: Sentiment Analysis

[0898] The server uses an emotion engine to analyze the user's emotions and recognize their emotional state from audio and video data. The input is the user's audio and video data, and the output is the evaluation result of their emotional state. Specifically, the emotion engine uses voice recognition and facial expression recognition algorithms to analyze the user's emotional state.

[0899] Step 7: Generate improvement suggestions and emergency alerts

[0900] The server generates and notifies improvement proposals and emergency alerts based on the analysis results. The input is the analysis results from steps 5 and 6, and the output is improvement proposals and emergency alerts. Specifically, the server generates optimal improvement proposals based on the analysis results and notifies the user device, such as a smartphone or tablet.

[0901] Step 8: Gather feedback and fine-tune the system

[0902] The user reviews the suggestions and alerts and provides their feedback to the server. The input is the user's feedback, and the output is the system's adjustment results. Specifically, the user enters their evaluation of the suggestions and alerts into a feedback form, and the server uses this information to improve the accuracy of the AI ​​model and analysis algorithms.

[0903] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0904] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0905] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0906] [Third embodiment]

[0907] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0908] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0909] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0910] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0911] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0912] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0913] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0914] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0915] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0916] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0917] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0918] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0919] This invention describes an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement proposals and emergency alerts. The operation of this system's program is explained below in natural language.

[0920] Data Collection Overview

[0921] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[0922] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[0923] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[0924] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[0925] Data Preprocessing Overview

[0926] 3. The server performs preprocessing on the collected data, specifically removing noise from the audio data, deleting unnecessary frames from the video data, and analyzing the text data.

[0927] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[0928] 4. The server converts the different types of data into a unified format.

[0929] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[0930] Overview of Data Integration and Model Training

[0931] 5. The server aggregates the pre-processed data to create a single aggregated dataset.

[0932] Example: The server stores cleaned video data, audio data, sensor data, and text data as a single dataset.

[0933] 6. The server trains the multimodal AI model using the integrated dataset.

[0934] Example: The server uses a deep learning framework to train models and improve its efficiency and safety decision-making capabilities.

[0935] Real-time analysis and improvement suggestions overview

[0936] 7. The server analyzes the collected real-time data and monitors the situation on site.

[0937] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[0938] 8. Based on the analysis results, the server identifies inefficient areas and potential dangers and generates appropriate improvement proposals.

[0939] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[0940] 9. The server will immediately alert the user if an imminent danger occurs.

[0941] Example: The server detects signs of a fire from video data and immediately sends a warning message to the administrator's terminal.

[0942] Notifications and Feedback Overview

[0943] 10. The server notifies the user of the generated improvement suggestions and emergency alerts.

[0944] Example: The server sends a message to the user's device saying, "Introducing a new work procedure will improve efficiency."

[0945] 11. The user checks the notified suggestions and alerts and takes the necessary action.

[0946] Example: A user reviews proposed work procedure changes and instructs field workers on the new procedures.

[0947] 12. Users provide feedback to the server on suggestions and alerts, helping to improve the system's accuracy.

[0948] Example: A user fills in a feedback form with a newly encountered problem and submits it to the server.

[0949] The above is a specific embodiment of the system of the present invention, which can significantly improve efficiency and safety in factories and construction sites.

[0950] The processing flow will be explained below.

[0951] Step 1:

[0952] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[0953] Step 2:

[0954] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[0955] Step 3:

[0956] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements. For example, it deletes frames in which human movement cannot be confirmed.

[0957] Step 4:

[0958] The server analyzes text data (accident reports, near-miss incident logs, work procedure manuals, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[0959] Step 5:

[0960] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[0961] Step 6:

[0962] The server then combines the pre-processed data to create a single integrated dataset, for example audio, video, temperature, and vibration data, which is stored in a cloud database and prepared for analysis.

[0963] Step 7:

[0964] The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[0965] Step 8:

[0966] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal conditions from real-time video and extracting abnormal sounds from audio data.

[0967] Step 9:

[0968] The server uses the analysis results to identify areas of inefficiency and potential dangers. For example, if a location where workers frequently stop or excessive vibration is detected, it will identify that as a potential danger.

[0969] Step 10:

[0970] The server generates improvement suggestions aimed at improving efficiency and safety, for example by changing the layout of machines or proposing new work procedures.

[0971] Step 11:

[0972] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[0973] Step 12:

[0974] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[0975] Step 13:

[0976] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[0977] Example 1

[0978] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0979] In factories and construction sites, real-time data collection and analysis is necessary to ensure both work efficiency and safety, but existing systems are difficult to adequately address this. There is also a need to integrate data from a variety of sensors and past data, and quickly generate appropriate improvement proposals and emergency alerts. Conventional technologies face challenges in unifying different data formats, removing noise, and improving the accuracy of real-time analysis.

[0980] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0981] In this invention, the server includes means for collecting data from input devices installed on-site, preprocessing means for removing noise and unnecessary data from the collected data, means for converting data of different formats into a unified format, means for integrating the preprocessed data to generate a dataset, means for training a multimodal model using the dataset, means for analyzing real-time data to determine efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting feedback from users to improve the accuracy of the system, thereby enabling significant improvements in work efficiency and safety at factories and construction sites.

[0982] A "server" is a central processing system that collects, processes, and analyzes data from various input devices installed in factories and construction sites, and notifies and presents the results to users.

[0983] "Input devices" are devices installed on-site to collect data, and include 3D cameras, microphones, temperature sensors, vibration sensors, etc.

[0984] The "data collection means" is a mechanism for continuously acquiring data from the input device and transmitting it to the server.

[0985] "Noise reduction means" refers to the algorithms and filtering processes used to remove background noise from collected audio data.

[0986] "Means for removing unnecessary data" refers to steps for eliminating frames in which no specific movement is detected or meaningless information from the video data.

[0987] "Preprocessing means" refers to a technique that includes processes such as noise removal and deletion of unnecessary data on collected raw data to make it ready for analysis.

[0988] "Means of converting into a unified format" refers to technology that unifies data collected in different formats into a consistent format, enabling integrated analysis.

[0989] "Dataset generation means" refers to the process of integrating preprocessed data and generating and saving a single integrated dataset.

[0990] A "multimodal model" is an artificial intelligence model that simultaneously analyzes different types of data (e.g., video, audio, text, sensor information) and learns their interrelationships.

[0991] "Model training methods" are techniques for training AI models using preprocessed and integrated datasets to improve their accuracy.

[0992] "Real-time data analysis means" refers to technology that instantly analyzes collected data and evaluates and monitors the current on-site situation.

[0993] "Means for determining efficiency and safety" refers to an algorithm that evaluates on-site work efficiency and safety based on the results of analyzing real-time data and makes specific judgments.

[0994] The "improvement proposal generation means" is a process for generating specific proposals for improving work efficiency and safety based on the results of data analysis.

[0995] The "emergency alert notification means" is a mechanism for immediately issuing a warning to the user when a serious danger is detected.

[0996] The "feedback collection means" is a mechanism for collecting user opinions and suggestions for improvement on a server and using them to improve and enhance the accuracy of the system.

[0997] The present invention relates to a system for optimizing the efficiency and safety of work in factories and construction sites. This system is composed of a server, an input device, and a user terminal.

[0998] The server collects data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. This data includes video data, audio data, temperature data, and vibration data. For example, the 3D camera streams video of the site in real time, and the microphone captures audio from the site. The temperature and vibration sensors collect their respective environmental data.

[0999] The server removes noise from the collected data and deletes unnecessary data. This preprocessing improves the accuracy of the analysis. Specifically, a voice recognition algorithm is used to remove noise from the audio data, and a motion detection algorithm is used to delete unnecessary frames from the video data. In addition, PDF-formatted work procedures and past accident reports are converted into text format using OCR technology.

[1000] The preprocessed data is then converted into a unified format by the server. This conversion integrates data from different formats into a consistent format. The converted data is then stored on the server as a unified dataset and analyzed using a multimodal model.

[1001] A multimodal model is an artificial intelligence model that simultaneously analyzes different types of data, optimally analyzing different modalities (video, audio, text, and sensor information). The server uses this multimodal model to determine efficiency and safety. For example, real-time video analysis can detect worker movements and identify unnatural movements. Audio analysis can also detect abnormal sounds on-site.

[1002] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and proposes relocation of the machines. In addition, if an emergency alert is required, the server immediately sends a warning message to the user. For example, it detects signs of a fire from video data and immediately sends a warning message to the user's device.

[1003] Users check improvement suggestions and emergency alerts notified by the server. Based on the notified information, they can change work procedures or instruct on-site workers on emergency responses. Users can also provide feedback on suggestions and alerts to the server, contributing to improving the accuracy of the system. For example, users can enter newly discovered problems or areas for improvement into a feedback form and send it to the server.

[1004] As a result, the efficiency and safety of work at factories and construction sites can be significantly improved. The system of the present invention can quickly and accurately monitor the situation at the site by analyzing real-time data, and provide appropriate improvement suggestions and emergency alerts.

[1005] An example of a prompt is, "Please analyze the current factory data and generate suggestions to improve safety." Using this prompt, the generative AI model can make appropriate analyses and suggestions.

[1006] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1007] Step 1:

[1008] The server collects data from input devices installed on-site. These include a 3D camera, microphone, temperature sensor, and vibration sensor. The server receives a real-time video stream from the 3D camera and simultaneously captures audio data from the microphone. Data from the temperature sensor and vibration sensor is also sent to the server. This allows for an immediate understanding of the situation on-site.

[1009] Input: 3D camera image data, microphone audio data, temperature sensor and vibration sensor data

[1010] Output: Collected raw data (video, audio, temperature, vibration)

[1011] Step 2:

[1012] The server performs preprocessing on the collected data. Specifically, it removes noise from the audio data and deletes unnecessary frames from the video data. To remove noise from the audio data, it uses a speech recognition algorithm to remove background noise. For the video data, it uses a motion detection algorithm to remove inactive frames. Furthermore, it converts work procedures and accident reports into text data using OCR.

[1013] Input: Collected raw data (video, audio, temperature, vibration), work procedures, accident reports

[1014] Output: Preprocessed data (noise-removed audio, video with unnecessary frames deleted, materials converted into text data)

[1015] Step 3:

[1016] The server converts the pre-processed data into a unified format, which aligns data from different formats into a consistent format, such as text data converted from PDF or sensor data converted to JSON format.

[1017] Input: Preprocessed data

[1018] Output: Dataset converted to a unified format

[1019] Step 4:

[1020] The server then combines the data in a unified format to generate a single comprehensive dataset. All data is integrated based on timestamps and stored in a database. This dataset is then used for model training and real-time analysis.

[1021] Input: Data converted into a unified format

[1022] Output: Unified dataset

[1023] Step 5:

[1024] The server trains a multimodal AI model using the integrated dataset. It uses a deep learning framework (e.g., TensorFlow) to analyze patterns in the data and improve its ability to make safety and efficiency decisions. This training allows the model to analyze different types of data simultaneously.

[1025] Input: Integrated dataset

[1026] Output: Trained multimodal AI model

[1027] Step 6:

[1028] The server analyzes the collected real-time data and monitors the situation on-site. Real-time video analysis detects worker movements and identifies unnatural movements. Audio analysis detects abnormal sounds.

[1029] Input: Real-time data (video, audio, temperature, vibration)

[1030] Output: Analysis results (on-site condition monitoring data)

[1031] Step 7:

[1032] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and suggests changing the machine's location. It also immediately sends a warning message to the user if an emergency alert is required. For example, it detects signs of a fire from video data and sends a warning message to the user's device.

[1033] Input: Analysis results

[1034] Output: Improvement suggestions and emergency alerts

[1035] Step 8:

[1036] The server notifies the generated improvement suggestions and emergency alerts to the user's device, which has the function to receive and display the suggestions and alerts.

[1037] Input: Improvement suggestions and emergency alerts

[1038] Output: Notification to user terminal

[1039] Step 9:

[1040] The user checks the notified improvement proposals and alerts and takes the necessary action. For example, the user checks the proposed changes to work procedures and instructs the on-site workers on the new procedures. In the case of an emergency alert, the user can quickly implement countermeasures.

[1041] Input: Improvement suggestions and emergency alerts

[1042] Output: User response

[1043] Step 10:

[1044] Users can provide feedback on suggestions and alerts to the server, helping to improve the system's accuracy. This feedback is reflected in the next model training.

[1045] Input: User feedback

[1046] Output: Feedback data to the server

[1047] The above is the specific processing flow of this system.

[1048] (Application example 1)

[1049] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1050] To maximize the efficiency and safety of on-site work, it is important to monitor the situation in real time and quickly and accurately identify inefficient or dangerous areas. However, conventional systems have difficulty integrating dispersed data and analyzing it in real time, resulting in low accuracy in improvement proposals and emergency alerts. Furthermore, a monitoring system using advanced data analysis technology was required to enable robots to work safely and effectively on-site.

[1051] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1052] In this invention, the server includes means for collecting information from input devices installed on-site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for training a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting, analyzing, and monitoring robot sensor data in real time. This makes it possible to optimize the efficiency and safety of on-site work, reducing risks and improving work safety, particularly at sites where robots work.

[1053] "Worksite" refers to a place where work is carried out, such as a factory or construction site.

[1054] "Input devices" refers to devices such as sensors and cameras installed on-site to collect data.

[1055] "Preprocessing" refers to the process of removing unnecessary parts from collected data and converting it into a format that is easy to analyze.

[1056] "Dataset" refers to a set of data that is collected, preprocessed, and integrated from multiple data sources.

[1057] A "model" is an algorithm that learns from collected and preprocessed data to accomplish a specific task.

[1058] "Real-time data" refers to data that is collected continuously along an ongoing timeline.

[1059] "Efficiency" refers to the degree to which work is carried out smoothly and without waste.

[1060] "Safety" refers to the degree to which work is carried out without hazards.

[1061] "Improvement proposals" refer to recommendations for improving efficiency or safety.

[1062] "Emergency Alert" means an emergency notification of immediate danger at a site.

[1063] "Robot" refers to a mechanical device that is programmed to perform tasks automatically.

[1064] "Sensor data" refers to data such as temperature, vibration, audio, and video collected by sensors.

[1065] "Monitoring" refers to the act of continuously observing the situation at a site and detecting any changes.

[1066] "Camera" refers to a device that captures images and collects them as digital data.

[1067] "Microphone" refers to a device for collecting sound.

[1068] "Temperature sensor" refers to a device that measures temperature and collects it as digital data.

[1069] A "vibration sensor" refers to a device that detects vibrations and collects them as digital data.

[1070] "Real-time" means happening at the present time, without delay.

[1071] "Notification" refers to the act of sending a message to convey information.

[1072] "Server" refers to a computer system for collecting, preprocessing, integrating, analyzing, and notifying data.

[1073] This invention is realized using an AI system to optimize the efficiency and safety of work on-site. The system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient areas and dangerous spots, and generates improvement suggestions and emergency alerts.

[1074] 1. Program Overview

[1075] The server collects data from input devices such as 3D cameras, microphones, temperature sensors, and vibration sensors. It then preprocesses and integrates the collected data to generate a dataset. This dataset is used to train an AI model that analyzes real-time on-site data. Based on the analysis results, it makes efficiency and safety judgments, generates improvement suggestions and emergency alerts, and notifies the user's device or the robot's display.

[1076] 2. Hardware and Software Used

[1077] 3D camera: Collects footage of the scene in real time.

[1078] Microphone: Collects audio data.

[1079] Temperature Sensor: Collects temperature data in the field.

[1080] Vibration Sensor: Collects vibration data.

[1081] Server: Collects, preprocesses, integrates, analyzes, trains models, and notifies data. In particular, it uses deep learning models using TensorFlow and Keras.

[1082] User device: Receives and displays improvement suggestions and emergency alerts.

[1083] Robot: Monitors on-site operations and presents the necessary information to users in an easy-to-understand format.

[1084] 3. Data processing and calculation

[1085] The server performs the following data processing.

[1086] Data collection: Collect data in real time from 3D cameras, microphones, temperature sensors, and vibration sensors.

[1087] Data preprocessing: Converting video data to grayscale, removing noise from audio data, detecting outliers in temperature and vibration data, etc.

[1088] Data integration: Integrate various pre-processed data to generate a single dataset.

[1089] Model training: Use TensorFlow and Keras to train a multimodal AI model based on the integrated dataset.

[1090] Real-time analysis: Trained models analyze collected real-time data to assess efficiency and safety.

[1091] Notification generation: Based on the analysis results, improvement suggestions and emergency alerts are automatically generated and sent to the user's device.

[1092] 4. Examples and prompts

[1093] A concrete example of this system in action is a robot in a factory. The robot operates safely by collecting data from 3D cameras and temperature sensors in the work area. The server analyzes this data in real time, identifies abnormal temperature rises or unnatural movements in the work area, and immediately sends an emergency alert to the manager's terminal.

[1094] Example prompt sentence:

[1095] "Please explain in detail how real-time data from 3D cameras and temperature sensors is analyzed to train AI models that determine efficiency and risk so that factory robots can continue to work safely."

[1096] By implementing this invention, the efficiency and safety of work in factories and construction sites can be significantly improved.

[1097] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1098] Step 1:

[1099] The server collects data from 3D cameras, microphones, temperature sensors, and vibration sensors installed on-site. Video data, audio data, temperature data, and vibration data obtained from input devices are input to the server. The collected data is in the form of a video stream for video, an audio stream for audio, and numerical data output from the sensors for temperature and vibration. This allows the real-time situation on-site to be accumulated on the server as digital data.

[1100] Step 2:

[1101] The server performs preprocessing on the collected data. This involves converting the video data to grayscale and removing noise from the audio data. It also detects and filters out abnormal values ​​in the temperature and vibration data. Specifically, OpenCV is used to convert each frame of the video data to grayscale, and an audio filtering algorithm is used to remove unwanted noise from the audio data. This results in clean data suitable for analysis.

[1102] Step 3:

[1103] The server then combines the pre-processed data into a single dataset, including video frames converted to grayscale, filtered audio data, and normalized temperature and vibration data, consolidating the various data types into a single, integrated format for further analysis.

[1104] Step 4:

[1105] The server trains an AI model using the integrated dataset. Specifically, it uses TensorFlow and Keras to train a deep learning model. The input data is a preprocessed and integrated dataset, and the output is a model for evaluating site efficiency and safety. This model learns based on historical data and real-time input data.

[1106] Step 5:

[1107] The server uses a trained AI model to analyze on-site data collected in real time. The input data is video data, audio data, temperature data, and vibration data collected in real time, and the model analyzes this data to evaluate efficiency and safety. The analysis results output a judgment that a particular task is inefficient or dangerous.

[1108] Step 6:

[1109] The server generates improvement suggestions and emergency alerts based on the analysis results. For example, if workers' movements in a specific area are unnatural, it will suggest changes to work procedures as an improvement suggestion. If temperature data is abnormally high, it will determine that there is a risk of fire and generate an emergency alert. These notifications are generated in a specific and actionable format.

[1110] Step 7:

[1111] The server notifies the generated improvement proposals and emergency alerts to the user's terminal or the robot's display. The notification content is displayed so that the user can immediately understand it and take measures. Specifically, the server sends email notifications using SMTP and displays warning messages on the robot's display.

[1112] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1113] This invention describes an AI system for optimizing the efficiency and safety of work in factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. Furthermore, this invention combines an emotion engine that recognizes the user's emotions, enabling it to provide appropriate feedback based on the user's emotional state.

[1114] Data Collection Overview

[1115] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. If there is a shortage, it also refers to backup data.

[1116] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[1117] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[1118] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[1119] Data Preprocessing Overview

[1120] 3. The server performs noise reduction on the collected audio data, specifically filtering out background noise and extracting only the target audio.

[1121] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[1122] 4. The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[1123] 5. The server converts different data formats into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[1124] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[1125] Overview of Data Integration and Model Training

[1126] 6. The server combines the pre-processed data to create a single combined data set, for example audio, video, temperature, and vibration data, and stores it in a cloud database, ready for analysis.

[1127] 7. The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1128] Example: The server uses past data to train an AI model to help determine future work efficiency and safety.

[1129] Introducing the Emotion Engine

[1130] 8. The server combines an emotion engine that recognizes the user's emotions and analyzes the user's emotional state. For example, it analyzes audio data acquired from a microphone and recognizes the user's emotional state (anger, sadness, joy, etc.).

[1131] Example: The server uses speech recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[1132] 9. The server analyzes the user's facial expressions from the video data and detects changes in emotions. For example, it analyzes video data acquired from a camera and determines whether the user is smiling or confused.

[1133] Example: The server uses a facial expression recognition algorithm to suggest appropriate actions to take if the user is feeling anxious.

[1134] Real-time analysis and improvement suggestions overview

[1135] 10. The server analyzes the collected real-time data and monitors the situation at the site. It detects worker movements and abnormal situations from real-time video and extracts abnormal sounds from audio data.

[1136] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[1137] 11. Based on the analysis results, the server identifies inefficiencies and potential dangers and generates appropriate improvement proposals, such as changing the layout of machines or proposing new work procedures.

[1138] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[1139] 12. The server will immediately send an alert to the user if an emergency occurs. If it detects signs of a fire, it will immediately send a warning message to the administrator's terminal.

[1140] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[1141] Notifications and Feedback Overview

[1142] 13. The server notifies the user of the generated improvement suggestions and emergency alerts. The server adjusts the content and timing of the feedback based on the user's emotional state.

[1143] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[1144] 14. The user checks the suggestions and alerts sent from the server and takes necessary action. For example, the user communicates the proposed new work procedures to the field workers and instructs them to carry them out immediately.

[1145] Example: A user reviews the proposed new work procedure and instructs field workers on the new procedure.

[1146] 15. Users contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[1147] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[1148] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[1149] The processing flow will be explained below.

[1150] Step 1:

[1151] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[1152] Step 2:

[1153] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[1154] Example: A server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone.

[1155] Step 3:

[1156] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements.

[1157] Example: The server removes still frames from video data and extracts only frames with movement.

[1158] Step 4:

[1159] The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases.

[1160] Example: The server analyzes accident reports and extracts important keywords such as "fall," "fire," and "machine malfunction."

[1161] Step 5:

[1162] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[1163] Example: A server converts PDF files into text format to create a dataset for multimodal AI.

[1164] Step 6:

[1165] The server aggregates the pre-processed data to create a single aggregated data set.

[1166] Example: The server combines cleaned audio, video, sensor, and text data to create a unified dataset.

[1167] Step 7:

[1168] The server uses the combined data set to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1169] Example: The server uses past data to train AI models to improve future work efficiency and safety.

[1170] Step 8:

[1171] The server uses an emotion engine that recognizes the user's emotions to analyze the user's voice data and recognize the user's emotional state (anger, sadness, joy, etc.).

[1172] Example: A server uses voice recognition and emotion analysis algorithms to determine whether a user is stressed.

[1173] Step 9:

[1174] The server analyzes the user's facial expressions from the video data and detects changes in emotions.

[1175] Example: The server uses a facial expression recognition algorithm to determine whether the user is smiling, confused, etc.

[1176] Step 10:

[1177] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal situations from real-time video and extracting abnormal sounds from audio data.

[1178] Example: The server uses real-time video and audio analysis to detect worker movements and identify unnatural movements.

[1179] Step 11:

[1180] Based on the analysis results, the server identifies inefficient areas and potential danger points and generates appropriate improvement proposals.

[1181] Example: The server analyzes the locations where workers frequently stop and suggests relocating those locations.

[1182] Step 12:

[1183] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[1184] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[1185] Step 13:

[1186] The server notifies the user of the generated improvement suggestions and emergency alerts, and adjusts the content and timing of the feedback based on the user's emotional state.

[1187] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[1188] Step 14:

[1189] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[1190] Example: A user reviews a proposed new work procedure and communicates the new procedure to workers in the field.

[1191] Step 15:

[1192] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[1193] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[1194] Example 2

[1195] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1196] Optimizing the efficiency and safety of on-site work requires real-time situational awareness and appropriate feedback. However, existing systems lack the ability to integrate and analyze collected data, making it difficult to respond quickly and accurately to specific issues. In addition, they are unable to provide feedback that takes into account the emotional state of workers, limiting further improvements in efficiency and safety.

[1197] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting information from input devices installed at the site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for learning a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for analyzing the emotional state of the user and providing appropriate feedback based on the emotional state. This makes it possible to analyze the situation at the site in real time, improve efficiency and safety, and provide appropriate feedback based on the emotional state of the worker.

[1198] "Worksite" refers to the location where work actually takes place, such as a factory or construction site.

[1199] "Input devices" refers to equipment including sensors, cameras, microphones, etc. that are installed on-site to collect information.

[1200] "Means for collecting information" refers to the method or technology for transmitting data obtained from the input device to the server.

[1201] "Preprocessing" refers to a series of steps taken to convert collected raw data into a form that is easier to analyze.

[1202] "Preprocessing means" refers to techniques and methods for performing preprocessing, such as noise removal, data transformation, and extraction of important information.

[1203] A "dataset" refers to a single set of data that is created by integrating multiple preprocessed data.

[1204] "Means for generating datasets" refers to methods and techniques for combining pre-processed data into a format that an AI model can learn from.

[1205] "Means of training the model" refers to the methods and techniques used to train the AI ​​model using the generated dataset.

[1206] "Real-time data" refers to the latest data collected from the field.

[1207] "Means of analysis" refers to the techniques and methods used to analyze collected data and determine its efficiency and safety.

[1208] "Decision result" refers to the conclusion or evaluation obtained based on the analyzed data.

[1209] "Improvement proposals" refer to specific proposals or methods for improving efficiency or safety.

[1210] An "urgent alert" is a warning issued when a danger or problem occurs that requires immediate action.

[1211] "Means for notification" refers to the methods and technologies for communicating generated improvement suggestions and emergency alerts to users.

[1212] "Emotional state" refers to a user's mental state or emotion, including, for example, joy, anger, sadness, and the like.

[1213] "Means for providing feedback" refers to methods and techniques for providing appropriate information or advice to a user based on the analyzed emotional state.

[1214] This invention is an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes the user's emotional state to provide appropriate feedback.

[1215] The system configuration is as follows:

[1216] Data collection

[1217] The server collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. For example, the server receives video streams from 3D cameras and simultaneously captures audio data from microphones. It also collects data on past accidents, near-miss incident logs, work procedures, completed volume data, and machine specification data.

[1218] Specific examples

[1219] The server simultaneously collects the video stream from the 3D camera and the audio data from the microphone.

[1220] The server retrieves past fall accident reports and work procedure manuals in PDF format and stores them in a database.

[1221] Data Preprocessing

[1222] The server performs noise reduction on the collected audio data. Specifically, it filters out background noise from the audio data and extracts only the target audio. It also removes frames from the video data in which no specific movement is detected.

[1223] In addition, the server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, and also converts data in different formats into a unified format.

[1224] Specific examples

[1225] The server uses a voice recognition algorithm to remove background noise from the voice data acquired from the microphone.

[1226] The server converts the PDF files into text format, generating a dataset that is easy to analyze.

[1227] Data integration and model training

[1228] The server combines the pre-processed data to create a single integrated data set, which is then used to train an AI model that uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1229] Specific examples

[1230] The server uses past data to train the AI ​​model, helping to determine future work efficiency and safety.

[1231] Introducing the Emotion Engine

[1232] The server combines an emotion engine to recognize the user's emotional state. For example, it analyzes audio data acquired from a microphone to recognize the user's emotional state (anger, sadness, joy, etc.). It also analyzes the user's facial expressions from video data to detect emotional fluctuations.

[1233] Specific examples

[1234] The server uses voice recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[1235] The server uses a facial expression recognition algorithm to suggest appropriate measures to address the user's anxiety.

[1236] Real-time analysis and improvement suggestions

[1237] The server analyzes the collected real-time data and monitors the situation at the site. Based on the analysis results, it identifies inefficiencies and potential dangers, generates improvement proposals, and immediately sends alerts if an emergency danger occurs, if necessary.

[1238] Specific examples

[1239] The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[1240] If the server detects a fire, it will send a real-time alert to the administrator's smartphone.

[1241] Notifications and Feedback

[1242] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the feedback based on the user's emotional state.

[1243] The user checks the suggestions and alerts sent from the server and takes necessary action. The user also provides feedback on the suggestions and alerts to the server, contributing to improving the accuracy of the system.

[1244] Specific examples

[1245] The server takes into consideration the user's emotional state and notifies them of improvement suggestions at a time when they are least stressed.

[1246] The user checks the proposed new work procedure and instructs the on-site workers on the new procedure.

[1247] The user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[1248] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[1249] Prompt Sentence Examples

[1250] "Explain how you can collect data from 3D cameras and microphones installed in a factory, denoise the audio data, and analyze the emotional state."

[1251] "Please explain in detail the steps to train the AI ​​model using past accident data and work procedures."

[1252] "How can we collect real-time data from the field to optimize safety and efficiency?"

[1253] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1254] Step 1:

[1255] The server collects information from input devices installed on-site. Specifically, it receives data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. Inputs include video data, audio data, temperature data, and vibration data. The collected raw data is then output.

[1256] Step 2:

[1257] The server then performs a noise reduction process on the collected audio data. Specifically, it uses a speech recognition algorithm to filter out background noise and extract only the target voice. The input includes raw audio data, and the output is clear audio data with noise removed.

[1258] Step 3:

[1259] The server performs motion detection on the collected video data and removes frames in which no specific motion is detected. Specifically, it uses a video analysis algorithm to remove frames without motion. The input includes raw video data, and the output is video data with motion.

[1260] Step 4:

[1261] The server analyzes text data (such as accident reports, near-miss incident logs, and work procedures) using natural language processing (NLP) technology to extract important keywords and phrases. Specifically, it uses an NLP algorithm to extract keywords such as "fall," "fire," and "machine malfunction" from the text. The input includes the text data, and the output is the extracted keywords and phrases.

[1262] Step 5:

[1263] The server converts data of different formats into a unified format and generates a single integrated dataset. Specifically, it converts PDF-formatted work procedures into text format and integrates various sensor data along a time axis. Inputs include text data, audio data, video data, and sensor data, and the output is an integrated dataset.

[1264] Step 6:

[1265] The server uses the combined dataset to train the AI ​​model, specifically, using a deep learning algorithm to train the model, with the combined dataset as input and the trained AI model as output.

[1266] Step 7:

[1267] The server uses a trained AI model to analyze the collected real-time data. Specifically, it uses the AI ​​model to determine the efficiency and safety of the site. The input includes real-time data, and the output is an evaluation result regarding efficiency and safety.

[1268] Step 8:

[1269] The server generates improvement proposals and emergency alerts based on the evaluation results and notifies the user. Specifically, it uses AI to evaluate the analysis results and generate proposals and alerts as necessary. The input includes the evaluation results, and the generated improvement proposals and alerts are obtained as output.

[1270] Step 9:

[1271] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes audio and video data to identify the user's emotional state. The input includes audio and video data, and the output is an analysis result related to the user's emotional state.

[1272] Step 10:

[1273] The server provides appropriate feedback based on the user's emotional state. Specifically, it adjusts the content and timing of the feedback based on the user's emotional state. The input includes the analysis results of the user's emotional state, and the output is an appropriate feedback message.

[1274] (Application example 2)

[1275] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1276] Maintaining high levels of efficiency and safety at the same time is difficult in on-site work, requiring immediate responses to on-site conditions. Furthermore, lack of appropriate feedback that takes into account the emotional state of workers can lead to employee stress and anxiety that negatively impacts efficiency and safety. The purpose of this invention is to solve these problems and provide an excellent working environment.

[1277] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1278] In this invention, the server includes means for collecting information from input devices installed at the site, means for pre-processing the collected information, and means for integrating the pre-processed information to generate a data set, thereby optimizing the efficiency and safety of the site and providing appropriate feedback that takes into account the emotional state of the user.

[1279] "Input devices installed on-site" are devices installed in factories and construction sites to collect various information, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[1280] "Means for preprocessing collected information" refers to a function for cleansing and processing initial data, such as removing noise from collected audio data and analyzing video data.

[1281] The "means for generating a dataset" is a function that integrates pre-processed data of various formats and converts them into a single consistent format.

[1282] "Means of learning the model" refers to the ability to train an AI model based on an integrated dataset using techniques such as deep learning.

[1283] The "means for determining efficiency and safety" is a function that uses a trained AI model to analyze data collected in real time and evaluate the efficiency and safety of work.

[1284] The "means for generating and notifying improvement proposals and emergency alerts" is a function that generates optimal improvement proposals based on the judgment results and notifies the user of emergency alerts as necessary.

[1285] The "means including an emotion engine" is a function that analyzes the user's voice and video data to recognize the user's emotional state and provides feedback based on the emotional state.

[1286] "Past accident data" refers to data containing detailed information about accidents that have occurred in the past.

[1287] A "near miss incident diary" is a record of near miss incidents and close calls that have occurred in the past.

[1288] A "work procedure manual" is a document that describes the procedures and methods for performing a specific task.

[1289] "Performance data" refers to data relating to the amount of work and results achieved on-site within a specific period of time.

[1290] "Machine specification data" refers to data that includes detailed technical specifications of machines and equipment used on-site.

[1291] "Multimodal AI" refers to AI technology that has the ability to comprehensively analyze data in different formats (such as audio, video, text, etc.).

[1292] This invention relates to an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site and performs real-time analysis to identify inefficient work areas and dangerous spots, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes and analyzes the user's emotional state to provide appropriate feedback to the user based on their emotional state.

[1293] System Overview

[1294] This system has the following main functions:

[1295] 1. Information gathering

[1296] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, and other devices installed on-site.

[1297] The collected data includes past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[1298] 2. Data Preprocessing

[1299] The server performs noise reduction processing on the collected audio data and deletes frames in which no specific movement is detected from the video data.

[1300] The text data is analyzed using natural language processing (NLP) techniques to extract important keywords and phrases.

[1301] Data of different formats is converted into a unified format.

[1302] 3. Data integration and model training

[1303] The server integrates the pre-processed data to create a single integrated dataset and stores it in a cloud database.

[1304] The integrated data set will be used to train AI models to improve efficiency and safety analysis capabilities.

[1305] 4. Introducing the Emotion Engine

[1306] To recognize the user's emotions, the server analyzes audio data obtained from a microphone and video data obtained from a camera to recognize the user's emotional state (anger, sadness, joy, etc.).

[1307] Providing appropriate feedback to the user based on the perceived emotional state.

[1308] 5. Real-time analysis and improvement suggestions

[1309] The server analyzes real-time data from the site and monitors work efficiency and safety.

[1310] Based on the analysis results, inefficient areas and potential danger areas are identified and appropriate improvement proposals are generated.

[1311] If an emergency occurs, an alert is sent to the user immediately.

[1312] 6. Notifications and Feedback

[1313] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the notifications based on the user's emotional state.

[1314] Users can check the suggestions and alerts, take necessary actions, and provide feedback on the suggestions and alerts to the server, thereby contributing to improving the accuracy of the system.

[1315] Hardware and software used

[1316] The system uses the following hardware and software:

[1317] Hardware: 3D camera, microphone, temperature sensor, vibration sensor, cloud database

[1318] Software: OpenCV, PyAudio, TensorFlow, HuggingFace Transformers, NLP tools

[1319] Specific examples

[1320] For example, if "fall accidents" occur frequently in a factory, the following prompts can be input into the generative AI model based on the collected data to obtain improvement suggestions:

[1321] Example prompt

[1322] Analyze all accident reports and sensor signal logs for tip-over accidents that have occurred over the past six months and generate specific proposals for improving safety. Improvement proposals may include real-time monitoring, changes to work procedures, and repositioning of machinery.

[1323] This makes the system a powerful tool for improving efficiency and safety on-site.

[1324] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1325] Step 1: Gather information

[1326] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. installed on-site. The input in this step is raw data from each sensor, and the output is an unprocessed data stream. Specifically, the server periodically polls data from each input device and transfers it to a cloud database.

[1327] Step 2: Data Preprocessing

[1328] The server performs noise reduction on the collected audio data and removes frames from the video data where no specific motion is detected. The input is the raw data collected in step 1, and the output is the pre-processed, clean data. Specifically, it applies a noise filter to the audio data and a frame removal algorithm to the video data.

[1329] Step 3: Data Integration

[1330] The server integrates preprocessed data in various formats to generate a single dataset. The input is preprocessed audio, video, temperature, vibration, and other data, and the output is a unified-format dataset. Specifically, the server links data in different formats using timestamps and stores them in a cloud database as a single integrated dataset.

[1331] Step 4: Model training

[1332] The server uses the dataset to train the AI ​​model. The input is the dataset generated in step 3, and the output is the trained AI model. Specifically, the server uses a deep learning algorithm to train the data. Here, libraries such as TensorFlow are used to train the model.

[1333] Step 5: Real-time analysis

[1334] The server analyzes real-time data from the site to determine efficiency and safety. The input is the data collected in real time and the trained model, and the output is the evaluation results of efficiency and safety. Specifically, the server uses the AI ​​model to analyze the real-time data and identify inefficient areas and potential dangerous areas.

[1335] Step 6: Sentiment Analysis

[1336] The server uses an emotion engine to analyze the user's emotions and recognize their emotional state from audio and video data. The input is the user's audio and video data, and the output is the evaluation result of their emotional state. Specifically, the emotion engine uses voice recognition and facial expression recognition algorithms to analyze the user's emotional state.

[1337] Step 7: Generate improvement suggestions and emergency alerts

[1338] The server generates and notifies improvement proposals and emergency alerts based on the analysis results. The input is the analysis results from steps 5 and 6, and the output is improvement proposals and emergency alerts. Specifically, the server generates optimal improvement proposals based on the analysis results and notifies the user device, such as a smartphone or tablet.

[1339] Step 8: Gather feedback and fine-tune the system

[1340] The user reviews the suggestions and alerts and provides their feedback to the server. The input is the user's feedback, and the output is the system's adjustment results. Specifically, the user enters their evaluation of the suggestions and alerts into a feedback form, and the server uses this information to improve the accuracy of the AI ​​model and analysis algorithms.

[1341] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1342] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1343] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1344] [Fourth embodiment]

[1345] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1346] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1347] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1348] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1349] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1350] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1351] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1352] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1353] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1354] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1355] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1356] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1357] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1358] This invention describes an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement proposals and emergency alerts. The operation of this system's program is explained below in natural language.

[1359] Data Collection Overview

[1360] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[1361] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[1362] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[1363] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[1364] Data Preprocessing Overview

[1365] 3. The server performs preprocessing on the collected data, specifically removing noise from the audio data, deleting unnecessary frames from the video data, and analyzing the text data.

[1366] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[1367] 4. The server converts the different types of data into a unified format.

[1368] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[1369] Overview of Data Integration and Model Training

[1370] 5. The server aggregates the pre-processed data to create a single aggregated dataset.

[1371] Example: The server stores cleaned video data, audio data, sensor data, and text data as a single dataset.

[1372] 6. The server trains the multimodal AI model using the integrated dataset.

[1373] Example: The server uses a deep learning framework to train models and improve its efficiency and safety decision-making capabilities.

[1374] Real-time analysis and improvement suggestions overview

[1375] 7. The server analyzes the collected real-time data and monitors the situation on site.

[1376] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[1377] 8. Based on the analysis results, the server identifies inefficient areas and potential dangers and generates appropriate improvement proposals.

[1378] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[1379] 9. The server will immediately alert the user if an imminent danger occurs.

[1380] Example: The server detects signs of a fire from video data and immediately sends a warning message to the administrator's terminal.

[1381] Notifications and Feedback Overview

[1382] 10. The server notifies the user of the generated improvement suggestions and emergency alerts.

[1383] Example: The server sends a message to the user's device saying, "Introducing a new work procedure will improve efficiency."

[1384] 11. The user checks the notified suggestions and alerts and takes the necessary action.

[1385] Example: A user reviews proposed work procedure changes and instructs field workers on the new procedures.

[1386] 12. Users provide feedback to the server on suggestions and alerts, helping to improve the system's accuracy.

[1387] Example: A user fills in a feedback form with a newly encountered problem and submits it to the server.

[1388] The above is a specific embodiment of the system of the present invention, which can significantly improve efficiency and safety in factories and construction sites.

[1389] The processing flow will be explained below.

[1390] Step 1:

[1391] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[1392] Step 2:

[1393] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[1394] Step 3:

[1395] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements. For example, it deletes frames in which human movement cannot be confirmed.

[1396] Step 4:

[1397] The server analyzes text data (accident reports, near-miss incident logs, work procedure manuals, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[1398] Step 5:

[1399] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[1400] Step 6:

[1401] The server then combines the pre-processed data to create a single integrated dataset, for example audio, video, temperature, and vibration data, which is stored in a cloud database and prepared for analysis.

[1402] Step 7:

[1403] The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1404] Step 8:

[1405] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal conditions from real-time video and extracting abnormal sounds from audio data.

[1406] Step 9:

[1407] The server uses the analysis results to identify areas of inefficiency and potential dangers. For example, if a location where workers frequently stop or excessive vibration is detected, it will identify that as a potential danger.

[1408] Step 10:

[1409] The server generates improvement suggestions aimed at improving efficiency and safety, for example by changing the layout of machines or proposing new work procedures.

[1410] Step 11:

[1411] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[1412] Step 12:

[1413] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[1414] Step 13:

[1415] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[1416] Example 1

[1417] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1418] In factories and construction sites, real-time data collection and analysis is necessary to ensure both work efficiency and safety, but existing systems are difficult to adequately address this. There is also a need to integrate data from a variety of sensors and past data, and quickly generate appropriate improvement proposals and emergency alerts. Conventional technologies face challenges in unifying different data formats, removing noise, and improving the accuracy of real-time analysis.

[1419] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1420] In this invention, the server includes means for collecting data from input devices installed on-site, preprocessing means for removing noise and unnecessary data from the collected data, means for converting data of different formats into a unified format, means for integrating the preprocessed data to generate a dataset, means for training a multimodal model using the dataset, means for analyzing real-time data to determine efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting feedback from users to improve the accuracy of the system, thereby enabling significant improvements in work efficiency and safety at factories and construction sites.

[1421] A "server" is a central processing system that collects, processes, and analyzes data from various input devices installed in factories and construction sites, and notifies and presents the results to users.

[1422] "Input devices" are devices installed on-site to collect data, and include 3D cameras, microphones, temperature sensors, vibration sensors, etc.

[1423] The "data collection means" is a mechanism for continuously acquiring data from the input device and transmitting it to the server.

[1424] "Noise reduction means" refers to the algorithms and filtering processes used to remove background noise from collected audio data.

[1425] "Means for removing unnecessary data" refers to steps for eliminating frames in which no specific movement is detected or meaningless information from the video data.

[1426] "Preprocessing means" refers to a technique that includes processes such as noise removal and deletion of unnecessary data on collected raw data to make it ready for analysis.

[1427] "Means of converting into a unified format" refers to technology that unifies data collected in different formats into a consistent format, enabling integrated analysis.

[1428] "Dataset generation means" refers to the process of integrating preprocessed data and generating and saving a single integrated dataset.

[1429] A "multimodal model" is an artificial intelligence model that simultaneously analyzes different types of data (e.g., video, audio, text, sensor information) and learns their interrelationships.

[1430] "Model training methods" are techniques for training AI models using preprocessed and integrated datasets to improve their accuracy.

[1431] "Real-time data analysis means" refers to technology that instantly analyzes collected data and evaluates and monitors the current on-site situation.

[1432] "Means for determining efficiency and safety" refers to an algorithm that evaluates on-site work efficiency and safety based on the results of analyzing real-time data and makes specific judgments.

[1433] The "improvement proposal generation means" is a process for generating specific proposals for improving work efficiency and safety based on the results of data analysis.

[1434] The "emergency alert notification means" is a mechanism for immediately issuing a warning to the user when a serious danger is detected.

[1435] The "feedback collection means" is a mechanism for collecting user opinions and suggestions for improvement on a server and using them to improve and enhance the accuracy of the system.

[1436] The present invention relates to a system for optimizing the efficiency and safety of work in factories and construction sites. This system is composed of a server, an input device, and a user terminal.

[1437] The server collects data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. This data includes video data, audio data, temperature data, and vibration data. For example, the 3D camera streams video of the site in real time, and the microphone captures audio from the site. The temperature and vibration sensors collect their respective environmental data.

[1438] The server removes noise from the collected data and deletes unnecessary data. This preprocessing improves the accuracy of the analysis. Specifically, a voice recognition algorithm is used to remove noise from the audio data, and a motion detection algorithm is used to delete unnecessary frames from the video data. In addition, PDF-formatted work procedures and past accident reports are converted into text format using OCR technology.

[1439] The preprocessed data is then converted into a unified format by the server. This conversion integrates data from different formats into a consistent format. The converted data is then stored on the server as a unified dataset and analyzed using a multimodal model.

[1440] A multimodal model is an artificial intelligence model that simultaneously analyzes different types of data, optimally analyzing different modalities (video, audio, text, and sensor information). The server uses this multimodal model to determine efficiency and safety. For example, real-time video analysis can detect worker movements and identify unnatural movements. Audio analysis can also detect abnormal sounds on-site.

[1441] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and proposes relocation of the machines. In addition, if an emergency alert is required, the server immediately sends a warning message to the user. For example, it detects signs of a fire from video data and immediately sends a warning message to the user's device.

[1442] Users check improvement suggestions and emergency alerts notified by the server. Based on the notified information, they can change work procedures or instruct on-site workers on emergency responses. Users can also provide feedback on suggestions and alerts to the server, contributing to improving the accuracy of the system. For example, users can enter newly discovered problems or areas for improvement into a feedback form and send it to the server.

[1443] As a result, the efficiency and safety of work at factories and construction sites can be significantly improved. The system of the present invention can quickly and accurately monitor the situation at the site by analyzing real-time data, and provide appropriate improvement suggestions and emergency alerts.

[1444] An example of a prompt is, "Please analyze the current factory data and generate suggestions to improve safety." Using this prompt, the generative AI model can make appropriate analyses and suggestions.

[1445] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1446] Step 1:

[1447] The server collects data from input devices installed on-site. These include a 3D camera, microphone, temperature sensor, and vibration sensor. The server receives a real-time video stream from the 3D camera and simultaneously captures audio data from the microphone. Data from the temperature sensor and vibration sensor is also sent to the server. This allows for an immediate understanding of the situation on-site.

[1448] Input: 3D camera image data, microphone audio data, temperature sensor and vibration sensor data

[1449] Output: Collected raw data (video, audio, temperature, vibration)

[1450] Step 2:

[1451] The server performs preprocessing on the collected data. Specifically, it removes noise from the audio data and deletes unnecessary frames from the video data. To remove noise from the audio data, it uses a speech recognition algorithm to remove background noise. For the video data, it uses a motion detection algorithm to remove inactive frames. Furthermore, it converts work procedures and accident reports into text data using OCR.

[1452] Input: Collected raw data (video, audio, temperature, vibration), work procedures, accident reports

[1453] Output: Preprocessed data (noise-removed audio, video with unnecessary frames deleted, materials converted into text data)

[1454] Step 3:

[1455] The server converts the pre-processed data into a unified format, which aligns data from different formats into a consistent format, such as text data converted from PDF or sensor data converted to JSON format.

[1456] Input: Preprocessed data

[1457] Output: Dataset converted to a unified format

[1458] Step 4:

[1459] The server then combines the data in a unified format to generate a single comprehensive dataset. All data is integrated based on timestamps and stored in a database. This dataset is then used for model training and real-time analysis.

[1460] Input: Data converted into a unified format

[1461] Output: Unified dataset

[1462] Step 5:

[1463] The server trains a multimodal AI model using the integrated dataset. It uses a deep learning framework (e.g., TensorFlow) to analyze patterns in the data and improve its ability to make safety and efficiency decisions. This training allows the model to analyze different types of data simultaneously.

[1464] Input: Integrated dataset

[1465] Output: Trained multimodal AI model

[1466] Step 6:

[1467] The server analyzes the collected real-time data and monitors the situation on-site. Real-time video analysis detects worker movements and identifies unnatural movements. Audio analysis detects abnormal sounds.

[1468] Input: Real-time data (video, audio, temperature, vibration)

[1469] Output: Analysis results (on-site condition monitoring data)

[1470] Step 7:

[1471] The server uses the analysis results to identify inefficient areas and potential dangers and generate improvement proposals. For example, it analyzes the behavior of workers who frequently stop around machines and suggests changing the machine's location. It also immediately sends a warning message to the user if an emergency alert is required. For example, it detects signs of a fire from video data and sends a warning message to the user's device.

[1472] Input: Analysis results

[1473] Output: Improvement suggestions and emergency alerts

[1474] Step 8:

[1475] The server notifies the generated improvement suggestions and emergency alerts to the user's device, which has the function to receive and display the suggestions and alerts.

[1476] Input: Improvement suggestions and emergency alerts

[1477] Output: Notification to user terminal

[1478] Step 9:

[1479] The user checks the notified improvement proposals and alerts and takes the necessary action. For example, the user checks the proposed changes to work procedures and instructs the on-site workers on the new procedures. In the case of an emergency alert, the user can quickly implement countermeasures.

[1480] Input: Improvement suggestions and emergency alerts

[1481] Output: User response

[1482] Step 10:

[1483] Users can provide feedback on suggestions and alerts to the server, helping to improve the system's accuracy. This feedback is reflected in the next model training.

[1484] Input: User feedback

[1485] Output: Feedback data to the server

[1486] The above is the specific processing flow of this system.

[1487] (Application example 1)

[1488] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1489] To maximize the efficiency and safety of on-site work, it is important to monitor the situation in real time and quickly and accurately identify inefficient or dangerous areas. However, conventional systems have difficulty integrating dispersed data and analyzing it in real time, resulting in low accuracy in improvement proposals and emergency alerts. Furthermore, a monitoring system using advanced data analysis technology was required to enable robots to work safely and effectively on-site.

[1490] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1491] In this invention, the server includes means for collecting information from input devices installed on-site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for training a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for collecting, analyzing, and monitoring robot sensor data in real time. This makes it possible to optimize the efficiency and safety of on-site work, reducing risks and improving work safety, particularly at sites where robots work.

[1492] "Worksite" refers to a place where work is carried out, such as a factory or construction site.

[1493] "Input devices" refers to devices such as sensors and cameras installed on-site to collect data.

[1494] "Preprocessing" refers to the process of removing unnecessary parts from collected data and converting it into a format that is easy to analyze.

[1495] "Dataset" refers to a set of data that is collected, preprocessed, and integrated from multiple data sources.

[1496] A "model" is an algorithm that learns from collected and preprocessed data to accomplish a specific task.

[1497] "Real-time data" refers to data that is collected continuously along an ongoing timeline.

[1498] "Efficiency" refers to the degree to which work is carried out smoothly and without waste.

[1499] "Safety" refers to the degree to which work is carried out without hazards.

[1500] "Improvement proposals" refer to recommendations for improving efficiency or safety.

[1501] "Emergency Alert" means an emergency notification of immediate danger at a site.

[1502] "Robot" refers to a mechanical device that is programmed to perform tasks automatically.

[1503] "Sensor data" refers to data such as temperature, vibration, audio, and video collected by sensors.

[1504] "Monitoring" refers to the act of continuously observing the situation at a site and detecting any changes.

[1505] "Camera" refers to a device that captures images and collects them as digital data.

[1506] "Microphone" refers to a device for collecting sound.

[1507] "Temperature sensor" refers to a device that measures temperature and collects it as digital data.

[1508] A "vibration sensor" refers to a device that detects vibrations and collects them as digital data.

[1509] "Real-time" means happening at the present time, without delay.

[1510] "Notification" refers to the act of sending a message to convey information.

[1511] "Server" refers to a computer system for collecting, preprocessing, integrating, analyzing, and notifying data.

[1512] This invention is realized using an AI system to optimize the efficiency and safety of work on-site. The system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient areas and dangerous spots, and generates improvement suggestions and emergency alerts.

[1513] 1. Program Overview

[1514] The server collects data from input devices such as 3D cameras, microphones, temperature sensors, and vibration sensors. It then preprocesses and integrates the collected data to generate a dataset. This dataset is used to train an AI model that analyzes real-time on-site data. Based on the analysis results, it makes efficiency and safety judgments, generates improvement suggestions and emergency alerts, and notifies the user's device or the robot's display.

[1515] 2. Hardware and Software Used

[1516] 3D camera: Collects footage of the scene in real time.

[1517] Microphone: Collects audio data.

[1518] Temperature Sensor: Collects temperature data in the field.

[1519] Vibration Sensor: Collects vibration data.

[1520] Server: Collects, preprocesses, integrates, analyzes, trains models, and notifies data. In particular, it uses deep learning models using TensorFlow and Keras.

[1521] User device: Receives and displays improvement suggestions and emergency alerts.

[1522] Robot: Monitors on-site operations and presents the necessary information to users in an easy-to-understand format.

[1523] 3. Data processing and calculation

[1524] The server performs the following data processing.

[1525] Data collection: Collect data in real time from 3D cameras, microphones, temperature sensors, and vibration sensors.

[1526] Data preprocessing: Converting video data to grayscale, removing noise from audio data, detecting outliers in temperature and vibration data, etc.

[1527] Data integration: Integrate various pre-processed data to generate a single dataset.

[1528] Model training: Use TensorFlow and Keras to train a multimodal AI model based on the integrated dataset.

[1529] Real-time analysis: Trained models analyze collected real-time data to assess efficiency and safety.

[1530] Notification generation: Based on the analysis results, improvement suggestions and emergency alerts are automatically generated and sent to the user's device.

[1531] 4. Examples and prompts

[1532] A concrete example of this system in action is a robot in a factory. The robot operates safely by collecting data from 3D cameras and temperature sensors in the work area. The server analyzes this data in real time, identifies abnormal temperature rises or unnatural movements in the work area, and immediately sends an emergency alert to the manager's terminal.

[1533] Example prompt sentence:

[1534] "Please explain in detail how real-time data from 3D cameras and temperature sensors is analyzed to train AI models that determine efficiency and risk so that factory robots can continue to work safely."

[1535] By implementing this invention, the efficiency and safety of work in factories and construction sites can be significantly improved.

[1536] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1537] Step 1:

[1538] The server collects data from 3D cameras, microphones, temperature sensors, and vibration sensors installed on-site. Video data, audio data, temperature data, and vibration data obtained from input devices are input to the server. The collected data is in the form of a video stream for video, an audio stream for audio, and numerical data output from the sensors for temperature and vibration. This allows the real-time situation on-site to be accumulated on the server as digital data.

[1539] Step 2:

[1540] The server performs preprocessing on the collected data. This involves converting the video data to grayscale and removing noise from the audio data. It also detects and filters out abnormal values ​​in the temperature and vibration data. Specifically, OpenCV is used to convert each frame of the video data to grayscale, and an audio filtering algorithm is used to remove unwanted noise from the audio data. This results in clean data suitable for analysis.

[1541] Step 3:

[1542] The server then combines the pre-processed data into a single dataset, including video frames converted to grayscale, filtered audio data, and normalized temperature and vibration data, consolidating the various data types into a single, integrated format for further analysis.

[1543] Step 4:

[1544] The server trains an AI model using the integrated dataset. Specifically, it uses TensorFlow and Keras to train a deep learning model. The input data is a preprocessed and integrated dataset, and the output is a model for evaluating site efficiency and safety. This model learns based on historical data and real-time input data.

[1545] Step 5:

[1546] The server uses a trained AI model to analyze on-site data collected in real time. The input data is video data, audio data, temperature data, and vibration data collected in real time, and the model analyzes this data to evaluate efficiency and safety. The analysis results output a judgment that a particular task is inefficient or dangerous.

[1547] Step 6:

[1548] The server generates improvement suggestions and emergency alerts based on the analysis results. For example, if workers' movements in a specific area are unnatural, it will suggest changes to work procedures as an improvement suggestion. If temperature data is abnormally high, it will determine that there is a risk of fire and generate an emergency alert. These notifications are generated in a specific and actionable format.

[1549] Step 7:

[1550] The server notifies the generated improvement proposals and emergency alerts to the user's terminal or the robot's display. The notification content is displayed so that the user can immediately understand it and take measures. Specifically, the server sends email notifications using SMTP and displays warning messages on the robot's display.

[1551] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1552] This invention describes an AI system for optimizing the efficiency and safety of work in factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. Furthermore, this invention combines an emotion engine that recognizes the user's emotions, enabling it to provide appropriate feedback based on the user's emotional state.

[1553] Data Collection Overview

[1554] 1. The server constantly collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. If there is a shortage, it also refers to backup data.

[1555] Example: A server receives a video stream from a 3D camera and simultaneously captures audio data from a microphone.

[1556] 2. The server retrieves all relevant information, such as past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[1557] Example: The server reads reports of past falls and work procedures (PDF format) and stores them in a database.

[1558] Data Preprocessing Overview

[1559] 3. The server performs noise reduction on the collected audio data, specifically filtering out background noise and extracting only the target audio.

[1560] For example: The server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone, or removes frames from video data in which specific motion is not detected.

[1561] 4. The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, such as "fall," "fire," and "machine malfunction."

[1562] 5. The server converts different data formats into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[1563] Example: A server converts PDF files into text format, generating a dataset that is easy for multimodal AI to analyze.

[1564] Overview of Data Integration and Model Training

[1565] 6. The server combines the pre-processed data to create a single combined data set, for example audio, video, temperature, and vibration data, and stores it in a cloud database, ready for analysis.

[1566] 7. The server uses the combined dataset to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1567] Example: The server uses past data to train an AI model to help determine future work efficiency and safety.

[1568] Introducing the Emotion Engine

[1569] 8. The server combines an emotion engine that recognizes the user's emotions and analyzes the user's emotional state. For example, it analyzes audio data acquired from a microphone and recognizes the user's emotional state (anger, sadness, joy, etc.).

[1570] Example: The server uses speech recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[1571] 9. The server analyzes the user's facial expressions from the video data and detects changes in emotions. For example, it analyzes video data acquired from a camera and determines whether the user is smiling or confused.

[1572] Example: The server uses a facial expression recognition algorithm to suggest appropriate actions to take if the user is feeling anxious.

[1573] Real-time analysis and improvement suggestions overview

[1574] 10. The server analyzes the collected real-time data and monitors the situation at the site. It detects worker movements and abnormal situations from real-time video and extracts abnormal sounds from audio data.

[1575] Example: The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[1576] 11. Based on the analysis results, the server identifies inefficiencies and potential dangers and generates appropriate improvement proposals, such as changing the layout of machines or proposing new work procedures.

[1577] Example: The server analyzes the behavior of workers who frequently stop around a machine and suggests changing the machine's layout.

[1578] 12. The server will immediately send an alert to the user if an emergency occurs. If it detects signs of a fire, it will immediately send a warning message to the administrator's terminal.

[1579] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[1580] Notifications and Feedback Overview

[1581] 13. The server notifies the user of the generated improvement suggestions and emergency alerts. The server adjusts the content and timing of the feedback based on the user's emotional state.

[1582] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[1583] 14. The user checks the suggestions and alerts sent from the server and takes necessary action. For example, the user communicates the proposed new work procedures to the field workers and instructs them to carry them out immediately.

[1584] Example: A user reviews the proposed new work procedure and instructs field workers on the new procedure.

[1585] 15. Users contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[1586] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[1587] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[1588] The processing flow will be explained below.

[1589] Step 1:

[1590] The server collects real-time data from 3D cameras, microphones, and various sensors (temperature sensors, vibration sensors, etc.) installed on-site. If there is a shortage, it also refers to backup data.

[1591] Step 2:

[1592] The server performs noise reduction on the collected audio data, specifically filtering out background noise from the audio data and extracting only the target audio.

[1593] Example: A server uses a speech recognition algorithm to remove background noise from audio data captured by a microphone.

[1594] Step 3:

[1595] The server removes unnecessary frames from the video data and selects the frames necessary for analyzing the worker's movements.

[1596] Example: The server removes still frames from video data and extracts only frames with movement.

[1597] Step 4:

[1598] The server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases.

[1599] Example: The server analyzes accident reports and extracts important keywords such as "fall," "fire," and "machine malfunction."

[1600] Step 5:

[1601] The server converts different types of data into a unified format, specifically converting PDF work instructions into text format and integrating various sensor data along a time axis.

[1602] Example: A server converts PDF files into text format to create a dataset for multimodal AI.

[1603] Step 6:

[1604] The server aggregates the pre-processed data to create a single aggregated data set.

[1605] Example: The server combines cleaned audio, video, sensor, and text data to create a unified dataset.

[1606] Step 7:

[1607] The server uses the combined data set to train an AI model, which uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1608] Example: The server uses past data to train AI models to improve future work efficiency and safety.

[1609] Step 8:

[1610] The server uses an emotion engine that recognizes the user's emotions to analyze the user's voice data and recognize the user's emotional state (anger, sadness, joy, etc.).

[1611] Example: A server uses voice recognition and emotion analysis algorithms to determine whether a user is stressed.

[1612] Step 9:

[1613] The server analyzes the user's facial expressions from the video data and detects changes in emotions.

[1614] Example: The server uses a facial expression recognition algorithm to determine whether the user is smiling, confused, etc.

[1615] Step 10:

[1616] The server analyzes the collected real-time data and monitors the situation at the site, detecting worker movements and abnormal situations from real-time video and extracting abnormal sounds from audio data.

[1617] Example: The server uses real-time video and audio analysis to detect worker movements and identify unnatural movements.

[1618] Step 11:

[1619] Based on the analysis results, the server identifies inefficient areas and potential danger points and generates appropriate improvement proposals.

[1620] Example: The server analyzes the locations where workers frequently stop and suggests relocating those locations.

[1621] Step 12:

[1622] The server immediately sends an alert to the user if an emergency occurs, and immediately sends a warning message to the administrator's terminal if it detects signs of a fire.

[1623] Example: If the server detects a fire, it sends a real-time alert to the administrator's smartphone.

[1624] Step 13:

[1625] The server notifies the user of the generated improvement suggestions and emergency alerts, and adjusts the content and timing of the feedback based on the user's emotional state.

[1626] Example: The server takes into account the user's emotional state and notifies them of improvement suggestions at times when they are least stressed.

[1627] Step 14:

[1628] The user checks the suggestions and alerts sent from the server and takes necessary action, such as communicating the proposed new work procedures to on-site workers and instructing them to carry them out immediately.

[1629] Example: A user reviews a proposed new work procedure and communicates the new procedure to workers in the field.

[1630] Step 15:

[1631] Users can contribute to improving the accuracy of the system by providing feedback on improvement suggestions and alerts to the server. For example, users can enter newly discovered problems and evaluation results into a feedback form and send it to the server.

[1632] Example: A user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[1633] Example 2

[1634] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1635] Optimizing the efficiency and safety of on-site work requires real-time situational awareness and appropriate feedback. However, existing systems lack the ability to integrate and analyze collected data, making it difficult to respond quickly and accurately to specific issues. In addition, they are unable to provide feedback that takes into account the emotional state of workers, limiting further improvements in efficiency and safety.

[1636] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting information from input devices installed at the site, means for preprocessing the collected information, means for integrating the preprocessed information to generate a dataset, means for learning a model using the dataset, means for analyzing real-time data from the site and determining efficiency and safety, means for generating and notifying improvement suggestions and emergency alerts based on the determination results, and means for analyzing the emotional state of the user and providing appropriate feedback based on the emotional state. This makes it possible to analyze the situation at the site in real time, improve efficiency and safety, and provide appropriate feedback based on the emotional state of the worker.

[1637] "Worksite" refers to the location where work actually takes place, such as a factory or construction site.

[1638] "Input devices" refers to equipment including sensors, cameras, microphones, etc. that are installed on-site to collect information.

[1639] "Means for collecting information" refers to the method or technology for transmitting data obtained from the input device to the server.

[1640] "Preprocessing" refers to a series of steps taken to convert collected raw data into a form that is easier to analyze.

[1641] "Preprocessing means" refers to techniques and methods for performing preprocessing, such as noise removal, data transformation, and extraction of important information.

[1642] A "dataset" refers to a single set of data that is created by integrating multiple preprocessed data.

[1643] "Means for generating datasets" refers to methods and techniques for combining pre-processed data into a format that an AI model can learn from.

[1644] "Means of training the model" refers to the methods and techniques used to train the AI ​​model using the generated dataset.

[1645] "Real-time data" refers to the latest data collected from the field.

[1646] "Means of analysis" refers to the techniques and methods used to analyze collected data and determine its efficiency and safety.

[1647] "Decision result" refers to the conclusion or evaluation obtained based on the analyzed data.

[1648] "Improvement proposals" refer to specific proposals or methods for improving efficiency or safety.

[1649] An "urgent alert" is a warning issued when a danger or problem occurs that requires immediate action.

[1650] "Means for notification" refers to the methods and technologies for communicating generated improvement suggestions and emergency alerts to users.

[1651] "Emotional state" refers to a user's mental state or emotion, including, for example, joy, anger, sadness, and the like.

[1652] "Means for providing feedback" refers to methods and techniques for providing appropriate information or advice to a user based on the analyzed emotional state.

[1653] This invention is an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site, analyzes it in real time, identifies inefficient work areas and dangerous areas, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes the user's emotional state to provide appropriate feedback.

[1654] The system configuration is as follows:

[1655] Data collection

[1656] The server collects real-time data from input devices installed on-site, such as 3D cameras, microphones, temperature sensors, and vibration sensors. For example, the server receives video streams from 3D cameras and simultaneously captures audio data from microphones. It also collects data on past accidents, near-miss incident logs, work procedures, completed volume data, and machine specification data.

[1657] Specific examples

[1658] The server simultaneously collects the video stream from the 3D camera and the audio data from the microphone.

[1659] The server retrieves past fall accident reports and work procedure manuals in PDF format and stores them in a database.

[1660] Data Preprocessing

[1661] The server performs noise reduction on the collected audio data. Specifically, it filters out background noise from the audio data and extracts only the target audio. It also removes frames from the video data in which no specific movement is detected.

[1662] In addition, the server analyzes text data (accident reports, near-miss incident logs, work procedures, etc.) using natural language processing (NLP) technology to extract important keywords and phrases, and also converts data in different formats into a unified format.

[1663] Specific examples

[1664] The server uses a voice recognition algorithm to remove background noise from the voice data acquired from the microphone.

[1665] The server converts the PDF files into text format, generating a dataset that is easy to analyze.

[1666] Data integration and model training

[1667] The server combines the pre-processed data to create a single integrated data set, which is then used to train an AI model that uses deep learning techniques to improve its efficiency and safety analysis capabilities.

[1668] Specific examples

[1669] The server uses past data to train the AI ​​model, helping to determine future work efficiency and safety.

[1670] Introducing the Emotion Engine

[1671] The server combines an emotion engine to recognize the user's emotional state. For example, it analyzes audio data acquired from a microphone to recognize the user's emotional state (anger, sadness, joy, etc.). It also analyzes the user's facial expressions from video data to detect emotional fluctuations.

[1672] Specific examples

[1673] The server uses voice recognition and emotion analysis algorithms to determine whether the user is feeling stressed and provides appropriate feedback.

[1674] The server uses a facial expression recognition algorithm to suggest appropriate measures to address the user's anxiety.

[1675] Real-time analysis and improvement suggestions

[1676] The server analyzes the collected real-time data and monitors the situation at the site. Based on the analysis results, it identifies inefficiencies and potential dangers, generates improvement proposals, and immediately sends alerts if an emergency danger occurs, if necessary.

[1677] Specific examples

[1678] The server uses real-time video analysis to detect worker movements and identify unnatural movements.

[1679] If the server detects a fire, it will send a real-time alert to the administrator's smartphone.

[1680] Notifications and Feedback

[1681] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the feedback based on the user's emotional state.

[1682] The user checks the suggestions and alerts sent from the server and takes necessary action. The user also provides feedback on the suggestions and alerts to the server, contributing to improving the accuracy of the system.

[1683] Specific examples

[1684] The server takes into consideration the user's emotional state and notifies them of improvement suggestions at a time when they are least stressed.

[1685] The user checks the proposed new work procedure and instructs the on-site workers on the new procedure.

[1686] The user evaluates the improvement suggestions generated by the system and provides feedback to the server.

[1687] The above is a specific embodiment of the system of the present invention. By combining this system with an emotion engine, it is possible to provide appropriate feedback based on the user's emotional state, further improving efficiency and safety on-site.

[1688] Prompt Sentence Examples

[1689] "Explain how you can collect data from 3D cameras and microphones installed in a factory, denoise the audio data, and analyze the emotional state."

[1690] "Please explain in detail the steps to train the AI ​​model using past accident data and work procedures."

[1691] "How can we collect real-time data from the field to optimize safety and efficiency?"

[1692] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1693] Step 1:

[1694] The server collects information from input devices installed on-site. Specifically, it receives data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. Inputs include video data, audio data, temperature data, and vibration data. The collected raw data is then output.

[1695] Step 2:

[1696] The server then performs a noise reduction process on the collected audio data. Specifically, it uses a speech recognition algorithm to filter out background noise and extract only the target voice. The input includes raw audio data, and the output is clear audio data with noise removed.

[1697] Step 3:

[1698] The server performs motion detection on the collected video data and removes frames in which no specific motion is detected. Specifically, it uses a video analysis algorithm to remove frames without motion. The input includes raw video data, and the output is video data with motion.

[1699] Step 4:

[1700] The server analyzes text data (such as accident reports, near-miss incident logs, and work procedures) using natural language processing (NLP) technology to extract important keywords and phrases. Specifically, it uses an NLP algorithm to extract keywords such as "fall," "fire," and "machine malfunction" from the text. The input includes the text data, and the output is the extracted keywords and phrases.

[1701] Step 5:

[1702] The server converts data of different formats into a unified format and generates a single integrated dataset. Specifically, it converts PDF-formatted work procedures into text format and integrates various sensor data along a time axis. Inputs include text data, audio data, video data, and sensor data, and the output is an integrated dataset.

[1703] Step 6:

[1704] The server uses the combined dataset to train the AI ​​model, specifically, using a deep learning algorithm to train the model, with the combined dataset as input and the trained AI model as output.

[1705] Step 7:

[1706] The server uses a trained AI model to analyze the collected real-time data. Specifically, it uses the AI ​​model to determine the efficiency and safety of the site. The input includes real-time data, and the output is an evaluation result regarding efficiency and safety.

[1707] Step 8:

[1708] The server generates improvement proposals and emergency alerts based on the evaluation results and notifies the user. Specifically, it uses AI to evaluate the analysis results and generate proposals and alerts as necessary. The input includes the evaluation results, and the generated improvement proposals and alerts are obtained as output.

[1709] Step 9:

[1710] The server uses an emotion engine to recognize the user's emotional state. Specifically, it analyzes audio and video data to identify the user's emotional state. The input includes audio and video data, and the output is an analysis result related to the user's emotional state.

[1711] Step 10:

[1712] The server provides appropriate feedback based on the user's emotional state. Specifically, it adjusts the content and timing of the feedback based on the user's emotional state. The input includes the analysis results of the user's emotional state, and the output is an appropriate feedback message.

[1713] (Application example 2)

[1714] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1715] Maintaining high levels of efficiency and safety at the same time is difficult in on-site work, requiring immediate responses to on-site conditions. Furthermore, lack of appropriate feedback that takes into account the emotional state of workers can lead to employee stress and anxiety that negatively impacts efficiency and safety. The purpose of this invention is to solve these problems and provide an excellent working environment.

[1716] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1717] In this invention, the server includes means for collecting information from input devices installed at the site, means for pre-processing the collected information, and means for integrating the pre-processed information to generate a data set, thereby optimizing the efficiency and safety of the site and providing appropriate feedback that takes into account the emotional state of the user.

[1718] "Input devices installed on-site" are devices installed in factories and construction sites to collect various information, such as 3D cameras, microphones, temperature sensors, and vibration sensors.

[1719] "Means for preprocessing collected information" refers to a function for cleansing and processing initial data, such as removing noise from collected audio data and analyzing video data.

[1720] The "means for generating a dataset" is a function that integrates pre-processed data of various formats and converts them into a single consistent format.

[1721] "Means of learning the model" refers to the ability to train an AI model based on an integrated dataset using techniques such as deep learning.

[1722] The "means for determining efficiency and safety" is a function that uses a trained AI model to analyze data collected in real time and evaluate the efficiency and safety of work.

[1723] The "means for generating and notifying improvement proposals and emergency alerts" is a function that generates optimal improvement proposals based on the judgment results and notifies the user of emergency alerts as necessary.

[1724] The "means including an emotion engine" is a function that analyzes the user's voice and video data to recognize the user's emotional state and provides feedback based on the emotional state.

[1725] "Past accident data" refers to data containing detailed information about accidents that have occurred in the past.

[1726] A "near miss incident diary" is a record of near miss incidents and close calls that have occurred in the past.

[1727] A "work procedure manual" is a document that describes the procedures and methods for performing a specific task.

[1728] "Performance data" refers to data relating to the amount of work and results achieved on-site within a specific period of time.

[1729] "Machine specification data" refers to data that includes detailed technical specifications of machines and equipment used on-site.

[1730] "Multimodal AI" refers to AI technology that has the ability to comprehensively analyze data in different formats (such as audio, video, text, etc.).

[1731] This invention relates to an AI system for optimizing the efficiency and safety of work at factories and construction sites. This system collects information from multiple input devices installed on-site and performs real-time analysis to identify inefficient work areas and dangerous spots, and generates improvement suggestions and emergency alerts. It also combines an emotion engine that recognizes and analyzes the user's emotional state to provide appropriate feedback to the user based on their emotional state.

[1732] System Overview

[1733] This system has the following main functions:

[1734] 1. Information gathering

[1735] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, and other devices installed on-site.

[1736] The collected data includes past accident data, near-miss incident logs, work procedures, completed work data, and machine specification data.

[1737] 2. Data Preprocessing

[1738] The server performs noise reduction processing on the collected audio data and deletes frames in which no specific movement is detected from the video data.

[1739] The text data is analyzed using natural language processing (NLP) techniques to extract important keywords and phrases.

[1740] Data of different formats is converted into a unified format.

[1741] 3. Data integration and model training

[1742] The server integrates the pre-processed data to create a single integrated dataset and stores it in a cloud database.

[1743] The integrated data set will be used to train AI models to improve efficiency and safety analysis capabilities.

[1744] 4. Introducing the Emotion Engine

[1745] To recognize the user's emotions, the server analyzes audio data obtained from a microphone and video data obtained from a camera to recognize the user's emotional state (anger, sadness, joy, etc.).

[1746] Providing appropriate feedback to the user based on the perceived emotional state.

[1747] 5. Real-time analysis and improvement suggestions

[1748] The server analyzes real-time data from the site and monitors work efficiency and safety.

[1749] Based on the analysis results, inefficient areas and potential danger areas are identified and appropriate improvement proposals are generated.

[1750] If an emergency occurs, an alert is sent to the user immediately.

[1751] 6. Notifications and Feedback

[1752] The server then sends the generated improvement suggestions and emergency alerts to the user's device, adjusting the content and timing of the notifications based on the user's emotional state.

[1753] Users can check the suggestions and alerts, take necessary actions, and provide feedback on the suggestions and alerts to the server, thereby contributing to improving the accuracy of the system.

[1754] Hardware and software used

[1755] The system uses the following hardware and software:

[1756] Hardware: 3D camera, microphone, temperature sensor, vibration sensor, cloud database

[1757] Software: OpenCV, PyAudio, TensorFlow, HuggingFace Transformers, NLP tools

[1758] Specific examples

[1759] For example, if "fall accidents" occur frequently in a factory, the following prompts can be input into the generative AI model based on the collected data to obtain improvement suggestions:

[1760] Example prompt

[1761] Analyze all accident reports and sensor signal logs for tip-over accidents that have occurred over the past six months and generate specific proposals for improving safety. Improvement proposals may include real-time monitoring, changes to work procedures, and repositioning of machinery.

[1762] This makes the system a powerful tool for improving efficiency and safety on-site.

[1763] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1764] Step 1: Gather information

[1765] The server collects data in real time from 3D cameras, microphones, temperature sensors, vibration sensors, etc. installed on-site. The input in this step is raw data from each sensor, and the output is an unprocessed data stream. Specifically, the server periodically polls data from each input device and transfers it to a cloud database.

[1766] Step 2: Data Preprocessing

[1767] The server performs noise reduction on the collected audio data and removes frames from the video data where no specific motion is detected. The input is the raw data collected in step 1, and the output is the pre-processed, clean data. Specifically, it applies a noise filter to the audio data and a frame removal algorithm to the video data.

[1768] Step 3: Data Integration

[1769] The server integrates preprocessed data in various formats to generate a single dataset. The input is preprocessed audio, video, temperature, vibration, and other data, and the output is a unified-format dataset. Specifically, the server links data in different formats using timestamps and stores them in a cloud database as a single integrated dataset.

[1770] Step 4: Model training

[1771] The server uses the dataset to train the AI ​​model. The input is the dataset generated in step 3, and the output is the trained AI model. Specifically, the server uses a deep learning algorithm to train the data. Here, libraries such as TensorFlow are used to train the model.

[1772] Step 5: Real-time analysis

[1773] The server analyzes real-time data from the site to determine efficiency and safety. The input is the data collected in real time and the trained model, and the output is the evaluation results of efficiency and safety. Specifically, the server uses the AI ​​model to analyze the real-time data and identify inefficient areas and potential dangerous areas.

[1774] Step 6: Sentiment Analysis

[1775] The server uses an emotion engine to analyze the user's emotions and recognize their emotional state from audio and video data. The input is the user's audio and video data, and the output is the evaluation result of their emotional state. Specifically, the emotion engine uses voice recognition and facial expression recognition algorithms to analyze the user's emotional state.

[1776] Step 7: Generate improvement suggestions and emergency alerts

[1777] The server generates and notifies improvement proposals and emergency alerts based on the analysis results. The input is the analysis results from steps 5 and 6, and the output is improvement proposals and emergency alerts. Specifically, the server generates optimal improvement proposals based on the analysis results and notifies the user device, such as a smartphone or tablet.

[1778] Step 8: Gather feedback and fine-tune the system

[1779] The user reviews the suggestions and alerts and provides their feedback to the server. The input is the user's feedback, and the output is the system's adjustment results. Specifically, the user enters their evaluation of the suggestions and alerts into a feedback form, and the server uses this information to improve the accuracy of the AI ​​model and analysis algorithms.

[1780] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1781] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1782] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1783] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1784] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1785] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1786] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1787] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1788] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1789] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1790] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1791] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1792] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is install...

Claims

1. means for collecting information from an input device installed on-site; means for pre-processing the collected information; means for combining the pre-processed information to generate a dataset; means for training a model using said dataset; A means of analyzing real-time data from the field to determine efficiency and safety; means for generating and notifying an improvement proposal and an emergency alert based on the determination result; A system including:

2. 2. The system according to claim 1, wherein the collected information includes past accident data, a log of near-miss events, work procedures, completed data, and machine specification data.

3. The system of claim 1, wherein the integrated data is analyzed using multimodal AI.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A