system
The system addresses deteriorating safety management by using real-time video processing and generative models to detect unsafe actions, notify personnel, and provide educational information, enhancing safety awareness and reducing accidents.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Labor shortages and lack of experience in the working environment have led to deteriorating safety management quality, increasing the risk of industrial accidents, necessitating a system that can monitor unsafe actions and conditions in real-time and provide efficient safety education without constant on-site monitoring.
A system that acquires real-time video information from on-site observation devices, processes it using a generative model to detect unsafe actions and conditions, evaluates the risk, and notifies safety management personnel, while also providing educational information to workers based on past databases.
Enables rapid detection and notification of unsafe conditions, supports effective safety education, and improves worker awareness by compensating for experience gaps, thereby reducing the risk of accidents.
Smart Images

Figure 2026073443000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Due to labor shortages and lack of experience in the working environment, the quality of safety management has deteriorated, and there is a current situation where the risk of industrial accidents has increased. To address this issue, a system that monitors the on-site situation in real time, reliably detects unsafe actions and conditions, and promptly notifies safety management personnel is necessary. Also, a method for efficiently conducting safety education for workers in a form that does not rely on constant on-site monitoring is required.
Means for Solving the Problems
[0005] This invention provides a system that acquires video information in real time from observation devices installed on-site and processes this video information to detect unsafe actions and conditions using a generative model. Furthermore, it includes means for evaluating the risk based on predetermined criteria based on the detection results and immediately notifying safety management personnel. In addition, this system enables effective safety education by creating improvement suggestions for the detected unsafe actions and conditions and providing them to workers as educational information. Moreover, by referring to a past database and extracting information on similar cases, it compensates for a lack of experience and enables prompt safety responses.
[0006] An "observation device" is a device installed to understand the situation on site and has the function of acquiring video information in real time.
[0007] "Visual information" refers to image and video data acquired by observation equipment, and represents the situation and actions at the site.
[0008] A "generative model" is an artificial intelligence algorithm that is trained based on past data and examples and used to detect unsafe behaviors and states.
[0009] "Unsafe actions" refer to actions that violate established safety standards in the work environment and increase the risk of accidents or injuries.
[0010] An "unsafe condition" refers to a situation in the work environment that contains risks due to unexpected malfunctions or inadequate maintenance, potentially leading to accidents.
[0011] "Risk assessment" is the process of evaluating the severity and frequency of detected unsafe actions or conditions and determining the degree of danger.
[0012] A "safety management officer" is a person responsible for maintaining safety standards in the workplace and ensuring the safety of workers.
[0013] "Educational information" refers to instructional data and materials provided to workers to teach them safe work practices. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The system of the present invention provides a complex set of functions combining observation devices, servers, and terminals to enhance safety management at the site. This makes it possible to detect unsafe operations and conditions in real time, and to perform risk assessment and notification.
[0036] First, observation equipment is installed at the work site, and cameras and sensors monitor the workers' movements and the work environment in real time. The server plays a central role in processing the video information received from the observation equipment and detects unsafe situations using a generative model. The generative model is trained on data collected under various conditions and automatically identifies unsafe behaviors and conditions.
[0037] If the server detects unsafe operation or conditions, it performs a risk assessment and immediately notifies the safety manager (user) of the results. The notification is made via a terminal, which warns the user of the need for countermeasures and prompts a rapid response on-site. Specific notifications include video information and detailed information about the risks, and in some cases, use voice or vibration to convey urgency.
[0038] Furthermore, the server provides specific safety measures that workers should take based on the problems it detects. This information is provided to employees as educational material via their terminals, leading to improved safety awareness among workers. In addition, it supports problem-solving more effectively by referring to similar cases from past databases.
[0039] As a concrete example, if a worker at a construction site is not wearing a safety harness while working at height, the observation device will detect this behavior and send video data to the server. The server will use a generative model to determine that the behavior is unsafe and will notify the user via the terminal that "failure to wear a safety harness while working at height" has been detected. The user can then take immediate action by following the instructions displayed on the terminal and provide appropriate guidance to the worker.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The server receives video data in real time from observation equipment installed at the work site. The received data is sent to the server without interruption using streaming technology.
[0043] Step 2:
[0044] The server preprocesses the received video data. Specifically, it adjusts the resolution, removes noise, extracts necessary areas, and converts the data into a format that is easy for generative models to analyze.
[0045] Step 3:
[0046] The server applies a generative model to pre-processed video data to detect unsafe behaviors and conditions. Based on a trained dataset, the model recognizes behaviors and conditions and compares them against safety standards.
[0047] Step 4:
[0048] The server performs a risk assessment of detected unsafe behaviors and conditions. It refers to a database of past cases and calculates the severity of the risk based on correlations with similar cases.
[0049] Step 5:
[0050] Based on the evaluation results, the server generates a notification for the terminal. This notification includes details of unsafe behavior or conditions, risk assessment results, and recommended countermeasures.
[0051] Step 6:
[0052] The device receives notifications sent from the server and displays them on the user interface. If necessary, it alerts the user using sound or vibration.
[0053] Step 7:
[0054] The user checks the notifications displayed on the device and decides on the appropriate response at the site. They then quickly communicate instructions to the workers at the site and, if necessary, take direct action.
[0055] Step 8:
[0056] Based on the educational information provided by the terminal, users will explain countermeasures for unsafe behaviors and conditions to workers and implement continuous safety training to further prevent workplace accidents.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] There is a need for a safety management system that can quickly and accurately detect unsafe behaviors and conditions on-site and prompt responsible personnel to take immediate action. However, conventional methods make it difficult to perform precise real-time monitoring and immediate response, and may not be effective in improving workers' safety awareness or providing training. Therefore, the development of a system that integrates high-precision detection, notification, and training support is necessary.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for acquiring information in real time from monitoring devices installed on-site, means for processing the acquired information and detecting unsafe behavior or conditions using a generation algorithm, and means for evaluating the risk based on the detected unsafe behavior or conditions according to set criteria and notifying the administrator. This enables real-time detection of unsafe conditions and rapid notification of risks.
[0062] A "monitoring device" is a device installed on-site that acquires real-time information on workers' actions and the environment.
[0063] A "generative algorithm" is a computational method that learns from past data and automatically identifies unsafe behaviors and states.
[0064] A "manager" is a person responsible for on-site safety management, receiving risk information, and taking appropriate action.
[0065] "Hazard" refers to the potential risks that arise when work or the work environment at a site deviates from safety standards.
[0066] "Notification" is the procedure for informing administrators about detected unsafe situations and the countermeasures being taken.
[0067] As an embodiment of the present invention, a method for constructing a system to enhance safety management will be described. This system provides comprehensive safety management with monitoring devices, servers, and terminals as its main components.
[0068] Hardware and software usage
[0069] The monitoring system is a device that combines cameras and sensors to monitor the movements of workers and the surrounding environment in real time. The information obtained from these devices is transmitted to a server.
[0070] The server functions as the central processing unit. This server is equipped with a generative AI model, which has been trained on data collected under various historical circumstances. This model analyzes data from the field and identifies unsafe behaviors and conditions.
[0071] Information about detected dangerous conditions is transmitted from the server to the terminals. These terminals consist of mobile devices and computers, allowing administrators to monitor the situation on-site and plan countermeasures in real time.
[0072] Specific examples and prompt statements
[0073] For example, at a construction site, if a monitoring device detects that a worker is not wearing the necessary safety equipment while working at height, the server uses a generated AI model to recognize this as an unsafe situation. The terminal notifies the administrator that "safety equipment is not being worn while working at height," prompting immediate action. In this case, the terminal will also provide an alert sound or vibration as needed.
[0074] Examples of prompts used to make a generative AI model work include the following:
[0075] "Monitor the work site in real time and determine whether safety standards are being met."
[0076] "Detect unsafe behaviors on-site and assess the risk level."
[0077] As described above, the system of the present invention surpasses conventional methods, enabling faster and more effective safety management.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The server acquires information in real time from monitoring devices installed on-site. These monitoring devices use cameras and sensors to capture worker movements and surrounding environment information, and transmit this data to the server. This data is used as input for analysis by the generated AI model.
[0081] Step 2:
[0082] The server analyzes the received information using a generating AI model. This model is trained on historical data and executes algorithms to identify unsafe behaviors and states. Specifically, it performs image analysis and behavioral pattern recognition, and compares the results with a database. As a result, data identified as unsafe is generated as output.
[0083] Step 3:
[0084] Unsafe conditions detected by the server are evaluated by a risk assessment algorithm. The server quantifies the risk against established criteria, and the evaluation results become the input for the next step. This reveals that high-priority situations require a rapid response.
[0085] Step 4:
[0086] Based on the evaluation results, the server sends a notification to the administrator via the terminal. The notification clearly states the details of the detected condition and the need for immediate action depending on the urgency. Specifically, it uses alert sounds and vibration functions to draw the administrator's attention and displays video information on the terminal.
[0087] Step 5:
[0088] Based on notifications and countermeasures received from the terminal, the user provides guidance to on-site workers. The server displays specific improvement measures and training information on the terminal, and the user uses this information to provide feedback to the workers. This process aims to improve safety awareness. In addition, if there are similar past cases, information referenced from the database is also added.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] Safety management at the workplace is extremely important for protecting the lives and health of workers. However, conventional safety management systems have problems such as not being able to quickly detect and warn of unsafe conditions in real time, and workers having difficulty obtaining necessary safety information immediately. Therefore, there is a need for a safety management system that can effectively provide immediate warnings and education to workers.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for acquiring video information in real time from observation devices installed on site, means for processing the acquired video information and detecting unsafe actions or conditions using a generation model, and means for presenting the safety status to workers in real time using a visual device and issuing audio and visual warnings as necessary. This enables immediate assessment of safety conditions at the work site and rapid warning.
[0094] An "observation device" is a piece of equipment installed on-site to acquire video information in real time.
[0095] "Visual information" refers to visual data acquired from observation equipment via cameras and sensors.
[0096] A "generative model" is a machine learning algorithm used to process video information and recognize and detect unsafe actions or states.
[0097] "Risk assessment" is the process of determining the severity and urgency of a risk based on detected unsafe behaviors or conditions, using predetermined criteria.
[0098] A "safety management officer" is responsible for safety management at the work site, receiving risk assessment results, and taking appropriate action.
[0099] A "visual device" is a device worn by workers that is used to visually present video information or warnings.
[0100] "Audio and visual warnings" are notification methods using sound and images that are used to inform workers of unsafe conditions.
[0101] The system implementing the present invention comprises an observation device, a server, and a visual device to enhance safety management at the work site. The observation device is installed at the site and monitors the worker's movements and the work environment in real time using cameras and sensors. Real-time video information transmitted from the observation device is received by the server, which uses a generative model based on this video information to detect unsafe movements and conditions.
[0102] The server analyzes video information in detail and can utilize Tensorflow® or PyTorch as its generation model. This allows it to automatically identify various unsafe conditions. If the server detects an unsafe condition, it performs a risk assessment based on that information and determines the severity of the risk according to predetermined criteria.
[0103] Next, the server immediately notifies the worker of the risk assessment results through a visual device. Smart glasses are used as the visual device, and by issuing audio and visual warnings, the worker can immediately understand the danger and take the necessary countermeasures. Specific examples of devices that serve as this interface include Google® Glass® and Microsoft® HoloLens®.
[0104] As a concrete example, consider a scenario where a worker accidentally enters an area outside of the designated control zone at a work site. In this case, the monitoring device detects the unsafe behavior, the server analyzes it, and immediately sends a warning to the smart glasses. This warning prompts the worker to quickly return to the safe area.
[0105] An example of a prompt might be: "Input frame images of the manufacturing area into the AI model and evaluate safety. If there are obstacles in the worker's path, issue a warning to the user. Provide suggestions for safety measures before proceeding to the next frame." This prompt provides important guidance for the AI model to support safety management in the workplace.
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The observation device acquires video information from the work site. This device uses a camera to capture the entire work environment in real time and transmits the acquired video as digital data to a server. The input is raw video from the camera, and the output is digitized image data.
[0109] Step 2:
[0110] The server receives video data from the observation device. The server divides the received image data into frames and inputs them into a generative AI model. In this process, the video data is compared with a database and features are extracted. The input is the divided image frames, and the output is a feature vector.
[0111] Step 3:
[0112] The server uses a generative AI model to analyze the input feature vector and detect unsafe behaviors and states. Based on the training data, the generative AI model identifies potential risks in the video. The input is a feature vector, and the output is the risk assessment result.
[0113] Step 4:
[0114] The server performs a risk assessment on detected risks based on established criteria and classifies the results in terms of importance and urgency. This assessment calculates the severity of the risk, and the appropriate action is determined according to the urgency. The input is the risk assessment result, and the output is the risk assessment score.
[0115] Step 5:
[0116] The server is connected to a visual device and communicates the assessed risk to the worker. Based on the risk score, the visual device presents the worker with audio and visual alerts. This process allows the worker to immediately take appropriate safety actions. The input is the risk assessment score, and the output is the alert notification.
[0117] Step 6:
[0118] The user (safety manager) reviews the information received via the visual device and issues instructions to the worker. Specifically, they instruct the worker to implement appropriate safety measures and provide ongoing safety training. The input is an alert notification, and the output is instructions and training information for the worker.
[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0120] This invention is a system that combines observation devices, a server, a terminal, and an emotion engine to enhance on-site safety management and raise safety awareness among workers. This system is characterized not only by its ability to detect unsafe actions and conditions in real time, but also by its ability to recognize user emotions and respond in a more personalized manner.
[0121] The system consists of observation devices installed on-site to capture video footage of workers' movements and surrounding conditions. A server receives this video data and uses a generative model to detect unsafe actions or conditions. At this time, a risk assessment is performed, and a notification based on the assessment results is sent to the terminal.
[0122] Furthermore, the addition of an emotion engine allows for the acquisition of information about the user's emotional state. The server analyzes this emotional data through the camera and microphone on the user's device and incorporates it into risk assessments and notification content. For example, if a user is experiencing high levels of stress, notifications can be made simpler and more specific to improve the efficiency of the response.
[0123] The device displays notifications and suggestions received from the server to the user. Because the presentation method is adjusted based on sentiment data, users receive information optimized according to their own emotional state.
[0124] For example, if a user is a safety manager at a work site and discovers a worker not wearing a safety harness while working at height, the server detects this risk and uses an emotion engine to determine that the user is in a state of anxiety. The terminal then displays countermeasures and past examples in a fast and easy-to-understand format. This process allows the user to give instructions quickly and calmly, which is expected to improve safety at the work site.
[0125] The following describes the processing flow.
[0126] Step 1:
[0127] The server receives real-time video data from observation devices installed on-site. This data includes worker movements and surrounding conditions and is used for safety assessments.
[0128] Step 2:
[0129] The server preprocesses the received video data and uses a generative model to detect unsafe behaviors and states. This model allows for the automatic identification of behaviors deemed dangerous.
[0130] Step 3:
[0131] When the server detects unsafe behavior or conditions, it performs a risk assessment. The assessment is based on pre-defined criteria and similar past cases, and determines the risk level of the detected event.
[0132] Step 4:
[0133] The server acquires emotional data from the device to recognize the user's emotions. It uses the device's camera and microphone to analyze the user's emotional state from their voice tone and facial expressions.
[0134] Step 5:
[0135] The server generates notifications and suggestions based on risk assessment results and the user's emotional state. When emotions are heightened, it strives to provide information that is more concise and easy to understand.
[0136] Step 6:
[0137] The device receives notifications sent from the server and displays them on the screen. These notifications include countermeasures against unsafe behavior and risk information that is sensitive to the user's feelings.
[0138] Step 7:
[0139] Users review notifications displayed on their devices and provide appropriate instructions to field workers based on the situation. Users can make calm and quick decisions based on information optimized for their emotions.
[0140] (Example 2)
[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0142] On-site safety management requires the detection and prompt response to unsafe actions and conditions. However, conventional technologies only detect and notify of hazards, failing to consider the emotional state of workers and managers, and sometimes resulting in inaccurate and swift responses. Furthermore, there is a challenge in effectively referencing past cases of dangerous situations to support rapid decision-making on-site.
[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0144] In this invention, the server includes means for acquiring image information in real time from a detection device installed on-site, means for processing the acquired image information and detecting dangerous actions or states using a generation artificial intelligence model, and means for analyzing the user's emotional information and adjusting the risk assessment and notification content based on the user's emotional state. This enables more rapid and appropriate management of dangerous situations in conjunction with emotional information, and allows for the provision of educational information that takes past cases into consideration.
[0145] A "detection device" is a device installed on-site to acquire image information in real time.
[0146] "Image information" refers to visual data that shows the actions of workers and the conditions of the environment.
[0147] A "generative artificial intelligence model" is a machine learning model trained to detect dangerous actions or conditions using acquired image information.
[0148] A "hazardous action or condition" is any action or surrounding condition that could pose a potential risk to workers or equipment.
[0149] "Risk assessment" is the process of estimating the degree of risk based on detected hazardous actions or conditions.
[0150] "User emotional information" refers to data that indicates the user's emotional state, and is obtained from information such as voice tone and facial expressions.
[0151] "Educational information" refers to instructional data provided to workers, including suggestions for improvement regarding dangerous actions or conditions.
[0152] A "safety management officer" is a person responsible for maintaining safety at the worksite and taking appropriate measures.
[0153] "Example information" refers to case study and example data related to specific dangerous actions or conditions, extracted from past information resources.
[0154] This invention is a system that utilizes a server, terminals, detection devices, and a generative AI model to enhance on-site safety management. The server processes real-time image information acquired from on-site detection devices and uses the generative AI model to detect dangerous actions or conditions. This system incorporates an emotion analysis engine that analyzes the user's emotional information and reflects it in notification content and risk assessment.
[0155] Specifically, the server utilizes hardware that receives data from cameras and sensors installed on-site. The received data is analyzed using generative artificial intelligence models that detect dangerous behaviors and states, employing machine learning frameworks such as TensorFlow and PyTorch. In addition, the sentiment analysis engine analyzes the user's emotional information using data collected from the user's device via camera and microphone.
[0156] If the user is a safety manager at the work site, the system will send a notification to the terminal that takes the user's emotional state into consideration when a specific hazard is detected. For example, if the hazard of not wearing safety equipment during work at height is detected and the user is determined to be in a state of agitation, a quick and clear suggestion will be displayed on the terminal. An example of such a prompt message would be: "Unsafe behavior observed at a work site at height. The user is in a state of agitation. Please briefly outline risk mitigation measures."
[0157] This allows users to respond quickly and accurately to hazards, thereby improving safety at the worksite.
[0158] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0159] Step 1:
[0160] The server receives image information in real time from on-site detection devices. This image information includes worker movements and environmental conditions. The received image data is preprocessed by noise reduction and conversion to an appropriate resolution, and then processed into a format suitable for analysis by AI models. The output is the processed image data.
[0161] Step 2:
[0162] The server passes processed image data as input to a generating artificial intelligence model to detect dangerous behaviors and states. This model is a deep learning model built using, for example, TensorFlow, and has already learned about dangerous behaviors and states. Data calculations include feature extraction and classification, and the degree of danger is evaluated. The output is a list of dangerous behaviors and states detected.
[0163] Step 3:
[0164] The server uses the camera and microphone on the user's device to collect emotional information in order to analyze the user's emotional state based on the risk detection results. This data includes facial recognition and voice tone analysis. For data processing, emotional characteristics are extracted from the video and audio and evaluated using an emotion analysis model. The output is the user's emotional state data.
[0165] Step 4:
[0166] The server integrates hazard detection results and emotional state data to generate notification content. This generation process creates prompt messages based on the urgency of the hazard and the user's emotional state. For example, it might say, "Unsafe behavior has been detected during work at height, and the user is in a state of anxiety. A concise immediate risk mitigation plan is presented." The output is a customized notification message.
[0167] Step 5:
[0168] The terminal displays notification messages sent from the server to the user. Here, specific actions are taken to adjust how notifications are displayed based on the user's emotional state. For example, for a user in an anxious state, larger fonts and high-contrast colors are used to make important information stand out. The output is an optimized information presentation for the user.
[0169] (Application Example 2)
[0170] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0171] Accidents caused by worker negligence or errors are frequent at work sites, leading to decreased productivity and personal injury. In this situation, there is a need for more efficient and accurate methods to ensure worker safety and raise safety awareness at the workplace. Furthermore, the lack of appropriate responses to workers' emotional states can hinder calm judgment.
[0172] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0173] In this invention, the server includes means for acquiring time-series data from observation devices installed in the work environment; means for processing the acquired time-series data and identifying dangerous behaviors or situations using a generative model; means for evaluating the danger based on defined criteria and notifying information providers based on the identified dangerous behaviors or situations; and means for collecting worker emotional information using an emotion analysis mechanism and adjusting the notification content based on the evaluation results. This enhances worker safety while enabling rapid on-site response.
[0174] "Observation devices installed in the work environment" are devices placed at the work site to grasp the conditions and activities of workers in that location in real time and collect them as time-series data.
[0175] "Time-series data" refers to data collected over time, allowing for sequential monitoring of the work environment and the status of workers.
[0176] A "generative model" is an algorithm that is trained on large amounts of data and used to identify specific patterns or abnormal behavior.
[0177] "Dangerous behavior or situation" refers to any action or condition within the work environment that increases the risk to workers or equipment, and includes those requiring safety management precautions.
[0178] "Defined criteria" are pre-established rules or conditions for assessing risks and determining appropriate responses.
[0179] An "information provider" is an individual or organization that is involved in safety management and is responsible for ensuring the safety of workers and equipment.
[0180] An "emotional analysis mechanism" is a technology or device used to analyze the emotional state of workers and acquire it as data.
[0181] "Adjusting notification content" refers to the process of changing or suppressing warning or guidance information depending on the emotional state of the worker.
[0182] The system that realizes this invention consists of an observation device, a server, a terminal, and an emotion analysis mechanism. First, the observation device is installed in the work environment and acquires time-series data of the worker and the work site. This data is transferred to the server via the network.
[0183] The server preprocesses data using image processing libraries such as OpenCV and identifies dangerous behaviors or situations by running generative AI models utilizing machine learning libraries such as TensorFlow. It also uses software such as Microsoft Azure Cognitive Services to analyze video and audio data of workers received from terminals and determine their emotional state. Based on the identified dangerous behaviors, the server evaluates the risk level and sends customized notifications to the terminals of information providers (workers and administrators who are users).
[0184] The terminal displays received notifications appropriately to the worker. The displayed content is adjusted according to the worker's emotional state; for example, a message encouraging calmness is added to a worker who is feeling anxious. In addition, support is provided to help workers make better decisions by referring to past case information.
[0185] As a concrete example, imagine a worker performing tasks next to a conveyor belt in a factory. The server performs emotion analysis, and if it detects that the worker is stressed, it displays a message on the smart glasses saying, "Assess the situation and proceed calmly to the next step."
[0186] Examples of prompts for the generating AI model include: "Propose a method to detect unsafe behavior in real time, analyze user sentiment data, and immediately display suggested responses." This prompt allows the system to generate appropriate information to support worker safety and efficient work.
[0187] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0188] Step 1:
[0189] The server receives time-series data from observation equipment installed in the work environment. This data includes camera images and sensor information. The server preprocesses the video data using OpenCV to remove noise and adjust the resolution. This allows for the acquisition of clean image data suitable for analysis.
[0190] Step 2:
[0191] The server analyzes the pre-processed video data using TensorFlow. It then runs a generative AI model to identify dangerous behaviors or situations in real time. In this step, the video data is given as input to the model, and the detected dangerous situations, their characteristics, and risk levels are returned as output.
[0192] Step 3:
[0193] The server uses Microsoft Azure Cognitive Services to perform emotion analysis. It takes audio and video data received from the worker's terminal as input and analyzes the worker's emotional state. The analysis results are output as an emotional state assessment (e.g., tension, anxiety, relaxation).
[0194] Step 4:
[0195] The server integrates identified risky behaviors and sentiment analysis results to create notifications for informants. These notifications are customized based on the risk assessment and emotional state of the generating AI model. For example, the notification might include a specific message such as, "Please calmly assess the situation and proceed cautiously with the next steps."
[0196] Step 5:
[0197] The terminal displays notifications received from the server to the user (worker). This display is done via smart glasses or a smartphone. By receiving these notifications, workers can obtain guidance for taking appropriate action.
[0198] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0199] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0200] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0201] [Second Embodiment]
[0202] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0203] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0204] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0205] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0206] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0207] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0208] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0209] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0210] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0211] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0212] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0213] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0214] The system of the present invention provides a complex set of functions combining observation devices, servers, and terminals to enhance safety management at the site. This makes it possible to detect unsafe operations and conditions in real time, and to perform risk assessment and notification.
[0215] First, observation equipment is installed at the work site, and cameras and sensors monitor the workers' movements and the work environment in real time. The server plays a central role in processing the video information received from the observation equipment and detects unsafe situations using a generative model. The generative model is trained on data collected under various conditions and automatically identifies unsafe behaviors and conditions.
[0216] If the server detects unsafe operation or conditions, it performs a risk assessment and immediately notifies the safety manager (user) of the results. The notification is made via a terminal, which warns the user of the need for countermeasures and prompts a rapid response on-site. Specific notifications include video information and detailed information about the risks, and in some cases, use voice or vibration to convey urgency.
[0217] Furthermore, the server provides specific safety measures that workers should take based on the problems it detects. This information is provided to employees as educational material via their terminals, leading to improved safety awareness among workers. In addition, it supports problem-solving more effectively by referring to similar cases from past databases.
[0218] As a concrete example, if a worker at a construction site is not wearing a safety harness while working at height, the observation device will detect this behavior and send video data to the server. The server will use a generative model to determine that the behavior is unsafe and will notify the user via the terminal that "failure to wear a safety harness while working at height" has been detected. The user can then take immediate action by following the instructions displayed on the terminal and provide appropriate guidance to the worker.
[0219] The following describes the processing flow.
[0220] Step 1:
[0221] The server receives video data in real time from observation equipment installed at the work site. The received data is sent to the server without interruption using streaming technology.
[0222] Step 2:
[0223] The server preprocesses the received video data. Specifically, it adjusts the resolution, removes noise, extracts necessary areas, and converts the data into a format that is easy for generative models to analyze.
[0224] Step 3:
[0225] The server applies a generative model to pre-processed video data to detect unsafe behaviors and conditions. Based on a trained dataset, the model recognizes behaviors and conditions and compares them against safety standards.
[0226] Step 4:
[0227] The server performs a risk assessment of detected unsafe behaviors and conditions. It refers to a database of past cases and calculates the severity of the risk based on correlations with similar cases.
[0228] Step 5:
[0229] Based on the evaluation results, the server generates a notification for the terminal. This notification includes details of unsafe behavior or conditions, risk assessment results, and recommended countermeasures.
[0230] Step 6:
[0231] The device receives notifications sent from the server and displays them on the user interface. If necessary, it alerts the user using sound or vibration.
[0232] Step 7:
[0233] The user checks the notifications displayed on the device and decides on the appropriate response at the site. They then quickly communicate instructions to the workers at the site and, if necessary, take direct action.
[0234] Step 8:
[0235] Based on the educational information provided by the terminal, users will explain countermeasures for unsafe behaviors and conditions to workers and implement continuous safety training to further prevent workplace accidents.
[0236] (Example 1)
[0237] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0238] There is a need for a safety management system that can quickly and accurately detect unsafe behaviors and conditions on-site and prompt responsible personnel to take immediate action. However, conventional methods make it difficult to perform precise real-time monitoring and immediate response, and may not be effective in improving workers' safety awareness or providing training. Therefore, the development of a system that integrates high-precision detection, notification, and training support is necessary.
[0239] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0240] In this invention, the server includes means for acquiring information in real time from monitoring devices installed on-site, means for processing the acquired information and detecting unsafe behavior or conditions using a generation algorithm, and means for evaluating the risk based on the detected unsafe behavior or conditions according to set criteria and notifying the administrator. This enables real-time detection of unsafe conditions and rapid notification of risks.
[0241] A "monitoring device" is a device installed on-site that acquires real-time information on workers' actions and the environment.
[0242] A "generative algorithm" is a computational method that learns from past data and automatically identifies unsafe behaviors and states.
[0243] A "manager" is a person responsible for on-site safety management, receiving risk information, and taking appropriate action.
[0244] "Hazard" refers to the potential risks that arise when work or the work environment at a site deviates from safety standards.
[0245] "Notification" is the procedure for informing administrators about detected unsafe situations and the countermeasures being taken.
[0246] As an embodiment of the present invention, a method for constructing a system to enhance safety management will be described. This system provides comprehensive safety management with monitoring devices, servers, and terminals as its main components.
[0247] Hardware and software usage
[0248] The monitoring system is a device that combines cameras and sensors to monitor the movements of workers and the surrounding environment in real time. The information obtained from these devices is transmitted to a server.
[0249] The server functions as the central processing unit. This server is equipped with a generative AI model, which has been trained on data collected under various historical circumstances. This model analyzes data from the field and identifies unsafe behaviors and conditions.
[0250] Information about detected dangerous conditions is transmitted from the server to the terminals. These terminals consist of mobile devices and computers, allowing administrators to monitor the situation on-site and plan countermeasures in real time.
[0251] Specific examples and prompt statements
[0252] For example, at a construction site, if a monitoring device detects that a worker is not wearing the necessary safety equipment while working at height, the server uses a generated AI model to recognize this as an unsafe situation. The terminal notifies the administrator that "safety equipment is not being worn while working at height," prompting immediate action. In this case, the terminal will also provide an alert sound or vibration as needed.
[0253] Examples of prompts used to make a generative AI model work include the following:
[0254] "Monitor the work site in real time and determine whether safety standards are being met."
[0255] "Detect unsafe behaviors on-site and assess the risk level."
[0256] As described above, the system of the present invention surpasses conventional methods, enabling faster and more effective safety management.
[0257] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0258] Step 1:
[0259] The server acquires information in real time from monitoring devices installed on-site. These monitoring devices use cameras and sensors to capture worker movements and surrounding environment information, and transmit this data to the server. This data is used as input for analysis by the generated AI model.
[0260] Step 2:
[0261] The server analyzes the received information using a generating AI model. This model is trained on historical data and executes algorithms to identify unsafe behaviors and states. Specifically, it performs image analysis and behavioral pattern recognition, and compares the results with a database. As a result, data identified as unsafe is generated as output.
[0262] Step 3:
[0263] Unsafe conditions detected by the server are evaluated by a risk assessment algorithm. The server quantifies the risk against established criteria, and the evaluation results become the input for the next step. This reveals that high-priority situations require a rapid response.
[0264] Step 4:
[0265] Based on the evaluation results, the server sends a notification to the administrator via the terminal. The notification clearly states the details of the detected condition and the need for immediate action depending on the urgency. Specifically, it uses alert sounds and vibration functions to draw the administrator's attention and displays video information on the terminal.
[0266] Step 5:
[0267] Based on notifications and countermeasures received from the terminal, the user provides guidance to on-site workers. The server displays specific improvement measures and training information on the terminal, and the user uses this information to provide feedback to the workers. This process aims to improve safety awareness. In addition, if there are similar past cases, information referenced from the database is also added.
[0268] (Application Example 1)
[0269] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0270] Safety management at the workplace is extremely important for protecting the lives and health of workers. However, conventional safety management systems have problems such as not being able to quickly detect and warn of unsafe conditions in real time, and workers having difficulty obtaining necessary safety information immediately. Therefore, there is a need for a safety management system that can effectively provide immediate warnings and education to workers.
[0271] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0272] In this invention, the server includes means for acquiring video information in real time from observation devices installed on site, means for processing the acquired video information and detecting unsafe actions or conditions using a generation model, and means for presenting the safety status to workers in real time using a visual device and issuing audio and visual warnings as necessary. This enables immediate assessment of safety conditions at the work site and rapid warning.
[0273] An "observation device" is a piece of equipment installed on-site to acquire video information in real time.
[0274] "Visual information" refers to visual data acquired from observation equipment via cameras and sensors.
[0275] A "generative model" is a machine learning algorithm used to process video information and recognize and detect unsafe actions or states.
[0276] "Risk assessment" is the process of determining the severity and urgency of a risk based on detected unsafe behaviors or conditions, using predetermined criteria.
[0277] A "safety management officer" is responsible for safety management at the work site, receiving risk assessment results, and taking appropriate action.
[0278] A "visual device" is a device worn by workers that is used to visually present video information or warnings.
[0279] "Audio and visual warnings" are notification methods using sound and images that are used to inform workers of unsafe conditions.
[0280] The system for implementing the present invention has a configuration including an observation device, a server, and a visual device in order to strengthen safety management at the work site. The observation device is installed on site and monitors the actions of workers and the working environment in real time by means of cameras and sensors. The real-time video information transmitted from the observation device is received by the server, and the server uses a generation model based on this video information to detect unsafe actions and states.
[0281] The server analyzes the video information in detail and can utilize TensorFlow or PyTorch as the generation model. Thereby, various unsafe states are automatically discriminated. When the server detects an unsafe state, it performs a risk assessment based on that information and determines the severity of the risk according to a predetermined criterion.
[0282] Next, the server immediately notifies the workers through the visual device of the risk assessment result. Smart glasses are used as the visual device, and by issuing audio and visual warnings, the workers can immediately grasp the danger and take necessary measures. Specific device examples of this interface include Google Glass and Microsoft HoloLens.
[0283] As a specific example, consider a scenario where a worker accidentally enters an area outside the managed range at the work site. At this time, the observation device detects the unsafe behavior, the server analyzes it, and immediately sends a warning to the smart glasses. Due to this warning, the worker is urged to quickly return to the safe range.
[0284] An example of a prompt sentence could be: "Input the frame image of the manufacturing area into the AI model and evaluate the safety. If there are obstacles in the worker's movement line, issue a warning to the user. Provide suggestions for safety measures before proceeding to the next frame." This prompt provides an important guideline for the AI model to assist in the safety management of the work site.
[0285] The flow of the specific process in Application Example 1 will be described using Fig. 12.
[0286] Step 1:
[0287] The observation device acquires video information of the work site. This device uses a camera to capture the entire work environment in real time and transmits the acquired video as digital data to the server. The input is the raw video from the camera, and the output is the digitized image data.
[0288] Step 2:
[0289] The server receives the video data received from the observation device. The server divides the received image data frame by frame and inputs it into the generated AI model. In this process, the video data is compared with the database, and features are extracted. The input is the divided image frame, and the output is the feature vector.
[0290] Step 3:
[0291] The server uses the generated AI model to analyze the input feature vector and detect unsafe operations or states. The generated AI model discriminates potential risks in the video based on the trained data. The input is the feature vector, and the output is the risk judgment result.
[0292] Step 4:
[0293] The server performs a risk assessment on the detected risks based on the criteria and classifies the results from the perspectives of importance and urgency. This assessment calculates the severity of the risk, and the processing is determined according to the urgency level. The input is the risk judgment result, and the output is the risk assessment score.
[0294] Step 5:
[0295] The server is connected to a visual device and communicates the assessed risk to the worker. Based on the risk score, the visual device presents the worker with audio and visual alerts. This process allows the worker to immediately take appropriate safety actions. The input is the risk assessment score, and the output is the alert notification.
[0296] Step 6:
[0297] The user (safety manager) reviews the information received via the visual device and issues instructions to the worker. Specifically, they instruct the worker to implement appropriate safety measures and provide ongoing safety training. The input is an alert notification, and the output is instructions and training information for the worker.
[0298] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0299] This invention is a system that combines observation devices, a server, a terminal, and an emotion engine to enhance on-site safety management and raise safety awareness among workers. This system is characterized not only by its ability to detect unsafe actions and conditions in real time, but also by its ability to recognize user emotions and respond in a more personalized manner.
[0300] The system consists of observation devices installed on-site to capture video footage of workers' movements and surrounding conditions. A server receives this video data and uses a generative model to detect unsafe actions or conditions. At this time, a risk assessment is performed, and a notification based on the assessment results is sent to the terminal.
[0301] Furthermore, the addition of an emotion engine allows for the acquisition of information about the user's emotional state. The server analyzes this emotional data through the camera and microphone on the user's device and incorporates it into risk assessments and notification content. For example, if a user is experiencing high levels of stress, notifications can be made simpler and more specific to improve the efficiency of the response.
[0302] The terminal shows the user the notifications and proposed content received from the server. Since the presentation method based on the emotional data is adjusted, the user can receive optimized information according to their emotional state.
[0303] As a specific example, if the user is a safety manager at the site and discovers a worker not wearing a safety belt during high-altitude work, the server detects this risk and determines through the emotion engine that the user is in a state of anxiety. Then, countermeasures and past cases are displayed on the terminal in a fast and easy-to-understand manner. Through this process, the user can give instructions quickly and calmly. As a result, an improvement in safety at the work site can be expected.
[0304] The following describes the processing flow.
[0305] Step 1:
[0306] The server receives real-time video data from the observation devices installed at the site. This data includes the actions of the workers and the surrounding situation, and is used for the evaluation of safety.
[0307] Step 2:
[0308] The server preprocesses the received video data and uses a generation model to detect unsafe actions and states. With this model, actions automatically determined to be dangerous can be identified.
[0309] Step 3:
[0310] When the server detects unsafe actions or states, it performs a risk assessment. The assessment is based on preset criteria and past similar cases to determine the risk level of the detected event.
[0311] [[ID=The server acquires emotional data from the device to recognize the user's emotions. It uses the device's camera and microphone to analyze the user's emotional state from their voice tone and facial expressions.
[0313] Step 5:
[0314] The server generates notifications and suggestions based on risk assessment results and the user's emotional state. When emotions are heightened, it strives to provide information that is more concise and easy to understand.
[0315] Step 6:
[0316] The device receives notifications sent from the server and displays them on the screen. These notifications include countermeasures against unsafe behavior and risk information that is sensitive to the user's feelings.
[0317] Step 7:
[0318] Users review notifications displayed on their devices and provide appropriate instructions to field workers based on the situation. Users can make calm and quick decisions based on information optimized for their emotions.
[0319] (Example 2)
[0320] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0321] On-site safety management requires the detection and prompt response to unsafe actions and conditions. However, conventional technologies only detect and notify of hazards, failing to consider the emotional state of workers and managers, and sometimes resulting in inaccurate and swift responses. Furthermore, there is a challenge in effectively referencing past cases of dangerous situations to support rapid decision-making on-site.
[0322] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0323] In this invention, the server includes means for acquiring image information in real time from a detection device installed on-site, means for processing the acquired image information and detecting dangerous actions or states using a generation artificial intelligence model, and means for analyzing the user's emotional information and adjusting the risk assessment and notification content based on the user's emotional state. This enables more rapid and appropriate management of dangerous situations in conjunction with emotional information, and allows for the provision of educational information that takes past cases into consideration.
[0324] A "detection device" is a device installed on-site to acquire image information in real time.
[0325] "Image information" refers to visual data that shows the actions of workers and the conditions of the environment.
[0326] A "generative artificial intelligence model" is a machine learning model trained to detect dangerous actions or conditions using acquired image information.
[0327] A "hazardous action or condition" is any action or surrounding condition that could pose a potential risk to workers or equipment.
[0328] "Risk assessment" is the process of estimating the degree of risk based on detected hazardous actions or conditions.
[0329] "User emotional information" refers to data that indicates the user's emotional state, and is obtained from information such as voice tone and facial expressions.
[0330] "Educational information" refers to instructional data provided to workers, including suggestions for improvement regarding dangerous actions or conditions.
[0331] A "safety management officer" is a person responsible for maintaining safety at the worksite and taking appropriate measures.
[0332] "Example information" refers to case study and example data related to specific dangerous actions or conditions, extracted from past information resources.
[0333] This invention is a system that utilizes a server, terminals, detection devices, and a generative AI model to enhance on-site safety management. The server processes real-time image information acquired from on-site detection devices and uses the generative AI model to detect dangerous actions or conditions. This system incorporates an emotion analysis engine that analyzes the user's emotional information and reflects it in notification content and risk assessment.
[0334] Specifically, the server utilizes hardware that receives data from cameras and sensors installed on-site. The received data is analyzed using generative artificial intelligence models that detect dangerous behaviors and states, employing machine learning frameworks such as TensorFlow and PyTorch. In addition, the sentiment analysis engine analyzes the user's emotional information using data collected from the user's device via camera and microphone.
[0335] If the user is a safety manager at the work site, the system will send a notification to the terminal that takes the user's emotional state into consideration when a specific hazard is detected. For example, if the hazard of not wearing safety equipment during work at height is detected and the user is determined to be in a state of agitation, a quick and clear suggestion will be displayed on the terminal. An example of such a prompt message would be: "Unsafe behavior observed at a work site at height. The user is in a state of agitation. Please briefly outline risk mitigation measures."
[0336] This allows users to respond quickly and accurately to hazards, thereby improving safety at the worksite.
[0337] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0338] Step 1:
[0339] The server receives image information in real time from on-site detection devices. This image information includes worker movements and environmental conditions. The received image data is preprocessed by noise reduction and conversion to an appropriate resolution, and then processed into a format suitable for analysis by AI models. The output is the processed image data.
[0340] Step 2:
[0341] The server passes processed image data as input to a generating artificial intelligence model to detect dangerous behaviors and states. This model is a deep learning model built using, for example, TensorFlow, and has already learned about dangerous behaviors and states. Data calculations include feature extraction and classification, and the degree of danger is evaluated. The output is a list of dangerous behaviors and states detected.
[0342] Step 3:
[0343] The server uses the camera and microphone on the user's device to collect emotional information in order to analyze the user's emotional state based on the risk detection results. This data includes facial recognition and voice tone analysis. For data processing, emotional characteristics are extracted from the video and audio and evaluated using an emotion analysis model. The output is the user's emotional state data.
[0344] Step 4:
[0345] The server integrates hazard detection results and emotional state data to generate notification content. This generation process creates prompt messages based on the urgency of the hazard and the user's emotional state. For example, it might say, "Unsafe behavior has been detected during work at height, and the user is in a state of anxiety. A concise immediate risk mitigation plan is presented." The output is a customized notification message.
[0346] Step 5:
[0347] The terminal displays notification messages sent from the server to the user. Here, specific actions are taken to adjust how notifications are displayed based on the user's emotional state. For example, for a user in an anxious state, larger fonts and high-contrast colors are used to make important information stand out. The output is an optimized information presentation for the user.
[0348] (Application Example 2)
[0349] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0350] Accidents caused by worker negligence or errors are frequent at work sites, leading to decreased productivity and personal injury. In this situation, there is a need for more efficient and accurate methods to ensure worker safety and raise safety awareness at the workplace. Furthermore, the lack of appropriate responses to workers' emotional states can hinder calm judgment.
[0351] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0352] In this invention, the server includes means for acquiring time-series data from observation devices installed in the work environment; means for processing the acquired time-series data and identifying dangerous behaviors or situations using a generative model; means for evaluating the danger based on defined criteria and notifying information providers based on the identified dangerous behaviors or situations; and means for collecting worker emotional information using an emotion analysis mechanism and adjusting the notification content based on the evaluation results. This enhances worker safety while enabling rapid on-site response.
[0353] "Observation devices installed in the work environment" are devices placed at the work site to grasp the conditions and activities of workers in that location in real time and collect them as time-series data.
[0354] "Time-series data" refers to data collected over time, allowing for sequential monitoring of the work environment and the status of workers.
[0355] A "generative model" is an algorithm that is trained on large amounts of data and used to identify specific patterns or abnormal behavior.
[0356] "Dangerous behavior or situation" refers to any action or condition within the work environment that increases the risk to workers or equipment, and includes those requiring safety management precautions.
[0357] "Defined criteria" are pre-established rules or conditions for assessing risks and determining appropriate responses.
[0358] An "information provider" is an individual or organization that is involved in safety management and is responsible for ensuring the safety of workers and equipment.
[0359] An "emotional analysis mechanism" is a technology or device used to analyze the emotional state of workers and acquire it as data.
[0360] "Adjusting notification content" refers to the process of changing or suppressing warning or guidance information depending on the emotional state of the worker.
[0361] The system that realizes this invention consists of an observation device, a server, a terminal, and an emotion analysis mechanism. First, the observation device is installed in the work environment and acquires time-series data of the worker and the work site. This data is transferred to the server via the network.
[0362] The server preprocesses data using image processing libraries such as OpenCV and identifies dangerous behaviors or situations by running generative AI models utilizing machine learning libraries such as TensorFlow. It also uses software such as Microsoft Azure Cognitive Services to analyze video and audio data received from terminals and determine the emotional state of the worker. Based on the identified dangerous behaviors, the server evaluates the risk level and sends a customized notification to the terminal of the information provider (the worker or manager who is the user).
[0363] The terminal displays received notifications appropriately to the worker. The displayed content is adjusted according to the worker's emotional state; for example, a message encouraging calmness is added to a worker who is feeling anxious. In addition, support is provided to help workers make better decisions by referring to past case information.
[0364] As a concrete example, imagine a worker performing tasks next to a conveyor belt in a factory. The server performs emotion analysis, and if it detects that the worker is stressed, it displays a message on the smart glasses saying, "Assess the situation and proceed calmly to the next step."
[0365] Examples of prompts for the generating AI model include: "Propose a method to detect unsafe behavior in real time, analyze user sentiment data, and immediately display suggested responses." This prompt allows the system to generate appropriate information to support worker safety and efficient work.
[0366] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0367] Step 1:
[0368] The server receives time-series data from observation equipment installed in the work environment. This data includes camera images and sensor information. The server preprocesses the video data using OpenCV to remove noise and adjust the resolution. This allows for the acquisition of clean image data suitable for analysis.
[0369] Step 2:
[0370] The server analyzes the pre-processed video data using TensorFlow. It then runs a generative AI model to identify dangerous behaviors or situations in real time. In this step, the video data is given as input to the model, and the detected dangerous situations, their characteristics, and risk levels are returned as output.
[0371] Step 3:
[0372] The server uses Microsoft Azure Cognitive Services to perform emotion analysis. It takes audio and video data received from the worker's terminal as input and analyzes the worker's emotional state. The analysis results are output as an emotional state assessment (e.g., tension, anxiety, relaxation).
[0373] Step 4:
[0374] The server integrates identified risky behaviors and sentiment analysis results to create notifications for informants. These notifications are customized based on the risk assessment and emotional state of the generating AI model. For example, the notification might include a specific message such as, "Please calmly assess the situation and proceed cautiously with the next steps."
[0375] Step 5:
[0376] The terminal displays notifications received from the server to the user (worker). This display is done via smart glasses or a smartphone. By receiving these notifications, workers can obtain guidance for taking appropriate action.
[0377] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0378] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0379] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0380] [Third Embodiment]
[0381] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0382] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0383] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0384] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0385] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0386] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0387] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0388] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0389] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0390] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0391] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0392] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0393] The system of the present invention provides a complex set of functions combining observation devices, servers, and terminals to enhance safety management at the site. This makes it possible to detect unsafe operations and conditions in real time, and to perform risk assessment and notification.
[0394] First, observation equipment is installed at the work site, and cameras and sensors monitor the workers' movements and the work environment in real time. The server plays a central role in processing the video information received from the observation equipment and detects unsafe situations using a generative model. The generative model is trained on data collected under various conditions and automatically identifies unsafe behaviors and conditions.
[0395] If the server detects unsafe operation or conditions, it performs a risk assessment and immediately notifies the safety manager (user) of the results. The notification is made via a terminal, which warns the user of the need for countermeasures and prompts a rapid response on-site. Specific notifications include video information and detailed information about the risks, and in some cases, use voice or vibration to convey urgency.
[0396] Furthermore, the server provides specific safety measures that workers should take based on the problems it detects. This information is provided to employees as educational material via their terminals, leading to improved safety awareness among workers. In addition, it supports problem-solving more effectively by referring to similar cases from past databases.
[0397] As a concrete example, if a worker at a construction site is not wearing a safety harness while working at height, the observation device will detect this behavior and send video data to the server. The server will use a generative model to determine that the behavior is unsafe and will notify the user via the terminal that "failure to wear a safety harness while working at height" has been detected. The user can then take immediate action by following the instructions displayed on the terminal and provide appropriate guidance to the worker.
[0398] The following describes the processing flow.
[0399] Step 1:
[0400] The server receives video data in real time from observation equipment installed at the work site. The received data is sent to the server without interruption using streaming technology.
[0401] Step 2:
[0402] The server preprocesses the received video data. Specifically, it adjusts the resolution, removes noise, extracts necessary areas, and converts the data into a format that is easy for generative models to analyze.
[0403] Step 3:
[0404] The server applies a generative model to pre-processed video data to detect unsafe behaviors and conditions. Based on a trained dataset, the model recognizes behaviors and conditions and compares them against safety standards.
[0405] Step 4:
[0406] The server performs a risk assessment of detected unsafe behaviors and conditions. It refers to a database of past cases and calculates the severity of the risk based on correlations with similar cases.
[0407] Step 5:
[0408] Based on the evaluation results, the server generates a notification for the terminal. This notification includes details of unsafe behavior or conditions, risk assessment results, and recommended countermeasures.
[0409] Step 6:
[0410] The device receives notifications sent from the server and displays them on the user interface. If necessary, it alerts the user using sound or vibration.
[0411] Step 7:
[0412] The user checks the notifications displayed on the device and decides on the appropriate response at the site. They then quickly communicate instructions to the workers at the site and, if necessary, take direct action.
[0413] Step 8:
[0414] Based on the educational information provided by the terminal, users will explain countermeasures for unsafe behaviors and conditions to workers and implement continuous safety training to further prevent workplace accidents.
[0415] (Example 1)
[0416] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0417] There is a need for a safety management system that can quickly and accurately detect unsafe behaviors and conditions on-site and prompt responsible personnel to take immediate action. However, conventional methods make it difficult to perform precise real-time monitoring and immediate response, and may not be effective in improving workers' safety awareness or providing training. Therefore, the development of a system that integrates high-precision detection, notification, and training support is necessary.
[0418] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0419] In this invention, the server includes means for acquiring information in real time from monitoring devices installed on-site, means for processing the acquired information and detecting unsafe behavior or conditions using a generation algorithm, and means for evaluating the risk based on the detected unsafe behavior or conditions according to set criteria and notifying the administrator. This enables real-time detection of unsafe conditions and rapid notification of risks.
[0420] A "monitoring device" is a device installed on-site that acquires real-time information on workers' actions and the environment.
[0421] A "generative algorithm" is a computational method that learns from past data and automatically identifies unsafe behaviors and states.
[0422] A "manager" is a person responsible for on-site safety management, receiving risk information, and taking appropriate action.
[0423] "Hazard" refers to the potential risks that arise when work or the work environment at a site deviates from safety standards.
[0424] "Notification" is the procedure for informing administrators about detected unsafe situations and the countermeasures being taken.
[0425] As an embodiment of the present invention, a method for constructing a system to enhance safety management will be described. This system provides comprehensive safety management with monitoring devices, servers, and terminals as its main components.
[0426] Hardware and software usage
[0427] The monitoring system is a device that combines cameras and sensors to monitor the movements of workers and the surrounding environment in real time. The information obtained from these devices is transmitted to a server.
[0428] The server functions as the central processing unit. This server is equipped with a generative AI model, which has been trained on data collected under various historical circumstances. This model analyzes data from the field and identifies unsafe behaviors and conditions.
[0429] Information about detected dangerous conditions is transmitted from the server to the terminals. These terminals consist of mobile devices and computers, allowing administrators to monitor the situation on-site and plan countermeasures in real time.
[0430] Specific examples and prompt statements
[0431] For example, at a construction site, if a monitoring device detects that a worker is not wearing the necessary safety equipment while working at height, the server uses a generated AI model to recognize this as an unsafe situation. The terminal notifies the administrator that "safety equipment is not being worn while working at height," prompting immediate action. In this case, the terminal will also provide an alert sound or vibration as needed.
[0432] Examples of prompts used to make a generative AI model work include the following:
[0433] "Monitor the work site in real time and determine whether safety standards are being met."
[0434] "Detect unsafe behaviors on-site and assess the risk level."
[0435] As described above, the system of the present invention surpasses conventional methods, enabling faster and more effective safety management.
[0436] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0437] Step 1:
[0438] The server acquires information in real time from monitoring devices installed on-site. These monitoring devices use cameras and sensors to capture worker movements and surrounding environment information, and transmit this data to the server. This data is used as input for analysis by the generated AI model.
[0439] Step 2:
[0440] The server analyzes the received information using a generating AI model. This model is trained on historical data and executes algorithms to identify unsafe behaviors and states. Specifically, it performs image analysis and behavioral pattern recognition, and compares the results with a database. As a result, data identified as unsafe is generated as output.
[0441] Step 3:
[0442] Unsafe conditions detected by the server are evaluated by a risk assessment algorithm. The server quantifies the risk against established criteria, and the evaluation results become the input for the next step. This reveals that high-priority situations require a rapid response.
[0443] Step 4:
[0444] Based on the evaluation results, the server sends a notification to the administrator via the terminal. The notification clearly states the details of the detected condition and the need for immediate action depending on the urgency. Specifically, it uses alert sounds and vibration functions to draw the administrator's attention and displays video information on the terminal.
[0445] Step 5:
[0446] Based on notifications and countermeasures received from the terminal, the user provides guidance to on-site workers. The server displays specific improvement measures and training information on the terminal, and the user uses this information to provide feedback to the workers. This process aims to improve safety awareness. In addition, if there are similar past cases, information referenced from the database is also added.
[0447] (Application Example 1)
[0448] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0449] Safety management at the workplace is extremely important for protecting the lives and health of workers. However, conventional safety management systems have problems such as not being able to quickly detect and warn of unsafe conditions in real time, and workers having difficulty obtaining necessary safety information immediately. Therefore, there is a need for a safety management system that can effectively provide immediate warnings and education to workers.
[0450] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0451] In this invention, the server includes means for acquiring video information in real time from observation devices installed on site, means for processing the acquired video information and detecting unsafe actions or conditions using a generation model, and means for presenting the safety status to workers in real time using a visual device and issuing audio and visual warnings as necessary. This enables immediate assessment of safety conditions at the work site and rapid warning.
[0452] An "observation device" is a piece of equipment installed on-site to acquire video information in real time.
[0453] "Visual information" refers to visual data acquired from observation equipment via cameras and sensors.
[0454] A "generative model" is a machine learning algorithm used to process video information and recognize and detect unsafe actions or states.
[0455] "Risk assessment" is the process of determining the severity and urgency of a risk based on detected unsafe behaviors or conditions, using predetermined criteria.
[0456] A "safety management officer" is responsible for safety management at the work site, receiving risk assessment results, and taking appropriate action.
[0457] A "visual device" is a device worn by workers that is used to visually present video information or warnings.
[0458] "Audio and visual warnings" are notification methods using sound and images that are used to inform workers of unsafe conditions.
[0459] The system implementing the present invention comprises an observation device, a server, and a visual device to enhance safety management at the work site. The observation device is installed at the site and monitors the worker's movements and the work environment in real time using cameras and sensors. Real-time video information transmitted from the observation device is received by the server, which uses a generative model based on this video information to detect unsafe movements and conditions.
[0460] The server analyzes video information in detail and can utilize TensorFlow or PyTorch as its generative model. This allows it to automatically identify various unsafe conditions. If the server detects an unsafe condition, it performs a risk assessment based on that information and determines the severity of the risk according to predetermined criteria.
[0461] Next, the server immediately notifies the worker of the risk assessment results through a visual device. Smart glasses are used as the visual device, and by issuing audio and visual warnings, the worker can immediately understand the hazard and take the necessary measures. Specific examples of devices that serve as this interface include Google Glass and Microsoft HoloLens.
[0462] As a concrete example, consider a scenario where a worker accidentally enters an area outside of the designated control zone at a work site. In this case, the monitoring device detects the unsafe behavior, the server analyzes it, and immediately sends a warning to the smart glasses. This warning prompts the worker to quickly return to the safe area.
[0463] An example of a prompt might be: "Input frame images of the manufacturing area into the AI model and evaluate safety. If there are obstacles in the worker's path, issue a warning to the user. Provide suggestions for safety measures before proceeding to the next frame." This prompt provides important guidance for the AI model to support safety management in the workplace.
[0464] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0465] Step 1:
[0466] The observation device acquires video information from the work site. This device uses a camera to capture the entire work environment in real time and transmits the acquired video as digital data to a server. The input is raw video from the camera, and the output is digitized image data.
[0467] Step 2:
[0468] The server receives video data from the observation device. The server divides the received image data into frames and inputs them into a generative AI model. In this process, the video data is compared with a database and features are extracted. The input is the divided image frames, and the output is a feature vector.
[0469] Step 3:
[0470] The server uses a generative AI model to analyze the input feature vector and detect unsafe behaviors and states. Based on the training data, the generative AI model identifies potential risks in the video. The input is a feature vector, and the output is the risk assessment result.
[0471] Step 4:
[0472] The server performs a risk assessment on detected risks based on established criteria and classifies the results in terms of importance and urgency. This assessment calculates the severity of the risk, and the appropriate action is determined according to the urgency. The input is the risk assessment result, and the output is the risk assessment score.
[0473] Step 5:
[0474] The server is connected to a visual device and communicates the assessed risk to the worker. Based on the risk score, the visual device presents the worker with audio and visual alerts. This process allows the worker to immediately take appropriate safety actions. The input is the risk assessment score, and the output is the alert notification.
[0475] Step 6:
[0476] The user (safety manager) reviews the information received via the visual device and issues instructions to the worker. Specifically, they instruct the worker to implement appropriate safety measures and provide ongoing safety training. The input is an alert notification, and the output is instructions and training information for the worker.
[0477] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0478] This invention is a system that combines observation devices, a server, a terminal, and an emotion engine to enhance on-site safety management and raise safety awareness among workers. This system is characterized not only by its ability to detect unsafe actions and conditions in real time, but also by its ability to recognize user emotions and respond in a more personalized manner.
[0479] The system consists of observation devices installed on-site to capture video footage of workers' movements and surrounding conditions. A server receives this video data and uses a generative model to detect unsafe actions or conditions. At this time, a risk assessment is performed, and a notification based on the assessment results is sent to the terminal.
[0480] Furthermore, the addition of an emotion engine allows for the acquisition of information about the user's emotional state. The server analyzes this emotional data through the camera and microphone on the user's device and incorporates it into risk assessments and notification content. For example, if a user is experiencing high levels of stress, notifications can be made simpler and more specific to improve the efficiency of the response.
[0481] The device displays notifications and suggestions received from the server to the user. Because the presentation method is adjusted based on sentiment data, users receive information optimized according to their own emotional state.
[0482] For example, if a user is a safety manager at a work site and discovers a worker not wearing a safety harness while working at height, the server detects this risk and uses an emotion engine to determine that the user is in a state of anxiety. The terminal then displays countermeasures and past examples in a fast and easy-to-understand format. This process allows the user to give instructions quickly and calmly, which is expected to improve safety at the work site.
[0483] The following describes the processing flow.
[0484] Step 1:
[0485] The server receives real-time video data from observation devices installed on-site. This data includes worker movements and surrounding conditions and is used for safety assessments.
[0486] Step 2:
[0487] The server preprocesses the received video data and uses a generative model to detect unsafe behaviors and states. This model allows for the automatic identification of behaviors deemed dangerous.
[0488] Step 3:
[0489] When the server detects unsafe behavior or conditions, it performs a risk assessment. The assessment is based on pre-defined criteria and similar past cases, and determines the risk level of the detected event.
[0490] Step 4:
[0491] The server acquires emotional data from the device to recognize the user's emotions. It uses the device's camera and microphone to analyze the user's emotional state from their voice tone and facial expressions.
[0492] Step 5:
[0493] The server generates notifications and suggestions based on risk assessment results and the user's emotional state. When emotions are heightened, it strives to provide information that is more concise and easy to understand.
[0494] Step 6:
[0495] The device receives notifications sent from the server and displays them on the screen. These notifications include countermeasures against unsafe behavior and risk information that is sensitive to the user's feelings.
[0496] Step 7:
[0497] Users review notifications displayed on their devices and provide appropriate instructions to field workers based on the situation. Users can make calm and quick decisions based on information optimized for their emotions.
[0498] (Example 2)
[0499] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0500] On-site safety management requires the detection and prompt response to unsafe actions and conditions. However, conventional technologies only detect and notify of hazards, failing to consider the emotional state of workers and managers, and sometimes resulting in inaccurate and swift responses. Furthermore, there is a challenge in effectively referencing past cases of dangerous situations to support rapid decision-making on-site.
[0501] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0502] In this invention, the server includes means for acquiring image information in real time from a detection device installed on-site, means for processing the acquired image information and detecting dangerous actions or states using a generation artificial intelligence model, and means for analyzing the user's emotional information and adjusting the risk assessment and notification content based on the user's emotional state. This enables more rapid and appropriate management of dangerous situations in conjunction with emotional information, and allows for the provision of educational information that takes past cases into consideration.
[0503] A "detection device" is a device installed on-site to acquire image information in real time.
[0504] "Image information" refers to visual data that shows the actions of workers and the conditions of the environment.
[0505] A "generative artificial intelligence model" is a machine learning model trained to detect dangerous actions or conditions using acquired image information.
[0506] A "hazardous action or condition" is any action or surrounding condition that could pose a potential risk to workers or equipment.
[0507] "Risk assessment" is the process of estimating the degree of risk based on detected hazardous actions or conditions.
[0508] "User emotional information" refers to data that indicates the user's emotional state, and is obtained from information such as voice tone and facial expressions.
[0509] "Educational information" refers to instructional data provided to workers, including suggestions for improvement regarding dangerous actions or conditions.
[0510] A "safety management officer" is a person responsible for maintaining safety at the worksite and taking appropriate measures.
[0511] "Example information" refers to case study and example data related to specific dangerous actions or conditions, extracted from past information resources.
[0512] This invention is a system that utilizes a server, terminals, detection devices, and a generative AI model to enhance on-site safety management. The server processes real-time image information acquired from on-site detection devices and uses the generative AI model to detect dangerous actions or conditions. This system incorporates an emotion analysis engine that analyzes the user's emotional information and reflects it in notification content and risk assessment.
[0513] Specifically, the server utilizes hardware that receives data from cameras and sensors installed on-site. The received data is analyzed using generative artificial intelligence models that detect dangerous behaviors and states, employing machine learning frameworks such as TensorFlow and PyTorch. In addition, the sentiment analysis engine analyzes the user's emotional information using data collected from the user's device via camera and microphone.
[0514] If the user is a safety manager at the work site, the system will send a notification to the terminal that takes the user's emotional state into consideration when a specific hazard is detected. For example, if the hazard of not wearing safety equipment during work at height is detected and the user is determined to be in a state of agitation, a quick and clear suggestion will be displayed on the terminal. An example of such a prompt message would be: "Unsafe behavior observed at a work site at height. The user is in a state of agitation. Please briefly outline risk mitigation measures."
[0515] This allows users to respond quickly and accurately to hazards, thereby improving safety at the worksite.
[0516] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0517] Step 1:
[0518] The server receives image information in real time from on-site detection devices. This image information includes worker movements and environmental conditions. The received image data is preprocessed by noise reduction and conversion to an appropriate resolution, and then processed into a format suitable for analysis by AI models. The output is the processed image data.
[0519] Step 2:
[0520] The server passes processed image data as input to a generating artificial intelligence model to detect dangerous behaviors and states. This model is a deep learning model built using, for example, TensorFlow, and has already learned about dangerous behaviors and states. Data calculations include feature extraction and classification, and the degree of danger is evaluated. The output is a list of dangerous behaviors and states detected.
[0521] Step 3:
[0522] The server uses the camera and microphone on the user's device to collect emotional information in order to analyze the user's emotional state based on the risk detection results. This data includes facial recognition and voice tone analysis. For data processing, emotional characteristics are extracted from the video and audio and evaluated using an emotion analysis model. The output is the user's emotional state data.
[0523] Step 4:
[0524] The server integrates hazard detection results and emotional state data to generate notification content. This generation process creates prompt messages based on the urgency of the hazard and the user's emotional state. For example, it might say, "Unsafe behavior has been detected during work at height, and the user is in a state of anxiety. A concise immediate risk mitigation plan is presented." The output is a customized notification message.
[0525] Step 5:
[0526] The terminal displays notification messages sent from the server to the user. Here, specific actions are taken to adjust how notifications are displayed based on the user's emotional state. For example, for a user in an anxious state, larger fonts and high-contrast colors are used to make important information stand out. The output is an optimized information presentation for the user.
[0527] (Application Example 2)
[0528] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0529] Accidents caused by worker negligence or errors are frequent at work sites, leading to decreased productivity and personal injury. In this situation, there is a need for more efficient and accurate methods to ensure worker safety and raise safety awareness at the workplace. Furthermore, the lack of appropriate responses to workers' emotional states can hinder calm judgment.
[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0531] In this invention, the server includes means for acquiring time-series data from observation devices installed in the work environment; means for processing the acquired time-series data and identifying dangerous behaviors or situations using a generative model; means for evaluating the danger based on defined criteria and notifying information providers based on the identified dangerous behaviors or situations; and means for collecting worker emotional information using an emotion analysis mechanism and adjusting the notification content based on the evaluation results. This enhances worker safety while enabling rapid on-site response.
[0532] "Observation devices installed in the work environment" are devices placed at the work site to grasp the conditions and activities of workers in that location in real time and collect them as time-series data.
[0533] "Time-series data" refers to data collected over time, allowing for sequential monitoring of the work environment and the status of workers.
[0534] A "generative model" is an algorithm that is trained on large amounts of data and used to identify specific patterns or abnormal behavior.
[0535] "Dangerous behavior or situation" refers to any action or condition within the work environment that increases the risk to workers or equipment, and includes those requiring safety management precautions.
[0536] "Defined criteria" are pre-established rules or conditions for assessing risks and determining appropriate responses.
[0537] An "information provider" is an individual or organization that is involved in safety management and is responsible for ensuring the safety of workers and equipment.
[0538] An "emotional analysis mechanism" is a technology or device used to analyze the emotional state of workers and acquire it as data.
[0539] "Adjusting notification content" refers to the process of changing or suppressing warning or guidance information depending on the emotional state of the worker.
[0540] The system that realizes this invention consists of an observation device, a server, a terminal, and an emotion analysis mechanism. First, the observation device is installed in the work environment and acquires time-series data of the worker and the work site. This data is transferred to the server via the network.
[0541] The server preprocesses data using image processing libraries such as OpenCV and identifies dangerous behaviors or situations by running generative AI models utilizing machine learning libraries such as TensorFlow. It also uses software such as Microsoft Azure Cognitive Services to analyze video and audio data received from terminals and determine the emotional state of the worker. Based on the identified dangerous behaviors, the server evaluates the risk level and sends a customized notification to the terminal of the information provider (the worker or manager who is the user).
[0542] The terminal displays received notifications appropriately to the worker. The displayed content is adjusted according to the worker's emotional state; for example, a message encouraging calmness is added to a worker who is feeling anxious. In addition, support is provided to help workers make better decisions by referring to past case information.
[0543] As a concrete example, imagine a worker performing tasks next to a conveyor belt in a factory. The server performs emotion analysis, and if it detects that the worker is stressed, it displays a message on the smart glasses saying, "Assess the situation and proceed calmly to the next step."
[0544] Examples of prompts for the generating AI model include: "Propose a method to detect unsafe behavior in real time, analyze user sentiment data, and immediately display suggested responses." This prompt allows the system to generate appropriate information to support worker safety and efficient work.
[0545] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0546] Step 1:
[0547] The server receives time-series data from observation equipment installed in the work environment. This data includes camera images and sensor information. The server preprocesses the video data using OpenCV to remove noise and adjust the resolution. This allows for the acquisition of clean image data suitable for analysis.
[0548] Step 2:
[0549] The server analyzes the pre-processed video data using TensorFlow. It then runs a generative AI model to identify dangerous behaviors or situations in real time. In this step, the video data is given as input to the model, and the detected dangerous situations, their characteristics, and risk levels are returned as output.
[0550] Step 3:
[0551] The server uses Microsoft Azure Cognitive Services to perform emotion analysis. It takes audio and video data received from the worker's terminal as input and analyzes the worker's emotional state. The analysis results are output as an emotional state assessment (e.g., tension, anxiety, relaxation).
[0552] Step 4:
[0553] The server integrates identified risky behaviors and sentiment analysis results to create notifications for informants. These notifications are customized based on the risk assessment and emotional state of the generating AI model. For example, the notification might include a specific message such as, "Please calmly assess the situation and proceed cautiously with the next steps."
[0554] Step 5:
[0555] The terminal displays notifications received from the server to the user (worker). This display is done via smart glasses or a smartphone. By receiving these notifications, workers can obtain guidance for taking appropriate action.
[0556] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0557] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0558] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0559] [Fourth Embodiment]
[0560] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0561] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0562] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0563] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0564] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0566] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0567] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0568] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0569] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0570] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0571] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0572] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0573] The system of the present invention provides a complex set of functions combining observation devices, servers, and terminals to enhance safety management at the site. This makes it possible to detect unsafe operations and conditions in real time, and to perform risk assessment and notification.
[0574] First, observation equipment is installed at the work site, and cameras and sensors monitor the workers' movements and the work environment in real time. The server plays a central role in processing the video information received from the observation equipment and detects unsafe situations using a generative model. The generative model is trained on data collected under various conditions and automatically identifies unsafe behaviors and conditions.
[0575] If the server detects unsafe operation or conditions, it performs a risk assessment and immediately notifies the safety manager (user) of the results. The notification is made via a terminal, which warns the user of the need for countermeasures and prompts a rapid response on-site. Specific notifications include video information and detailed information about the risks, and in some cases, use voice or vibration to convey urgency.
[0576] Furthermore, the server provides specific safety measures that workers should take based on the problems it detects. This information is provided to employees as educational material via their terminals, leading to improved safety awareness among workers. In addition, it supports problem-solving more effectively by referring to similar cases from past databases.
[0577] As a concrete example, if a worker at a construction site is not wearing a safety harness while working at height, the observation device will detect this behavior and send video data to the server. The server will use a generative model to determine that the behavior is unsafe and will notify the user via the terminal that "failure to wear a safety harness while working at height" has been detected. The user can then take immediate action by following the instructions displayed on the terminal and provide appropriate guidance to the worker.
[0578] The following describes the processing flow.
[0579] Step 1:
[0580] The server receives video data in real time from observation equipment installed at the work site. The received data is sent to the server without interruption using streaming technology.
[0581] Step 2:
[0582] The server preprocesses the received video data. Specifically, it adjusts the resolution, removes noise, extracts necessary areas, and converts the data into a format that is easy for generative models to analyze.
[0583] Step 3:
[0584] The server applies a generative model to pre-processed video data to detect unsafe behaviors and conditions. Based on a trained dataset, the model recognizes behaviors and conditions and compares them against safety standards.
[0585] Step 4:
[0586] The server performs a risk assessment of detected unsafe behaviors and conditions. It refers to a database of past cases and calculates the severity of the risk based on correlations with similar cases.
[0587] Step 5:
[0588] Based on the evaluation results, the server generates a notification for the terminal. This notification includes details of unsafe behavior or conditions, risk assessment results, and recommended countermeasures.
[0589] Step 6:
[0590] The device receives notifications sent from the server and displays them on the user interface. If necessary, it alerts the user using sound or vibration.
[0591] Step 7:
[0592] The user checks the notifications displayed on the device and decides on the appropriate response at the site. They then quickly communicate instructions to the workers at the site and, if necessary, take direct action.
[0593] Step 8:
[0594] Based on the educational information provided by the terminal, users will explain countermeasures for unsafe behaviors and conditions to workers and implement continuous safety training to further prevent workplace accidents.
[0595] (Example 1)
[0596] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0597] There is a need for a safety management system that can quickly and accurately detect unsafe behaviors and conditions on-site and prompt responsible personnel to take immediate action. However, conventional methods make it difficult to perform precise real-time monitoring and immediate response, and may not be effective in improving workers' safety awareness or providing training. Therefore, the development of a system that integrates high-precision detection, notification, and training support is necessary.
[0598] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0599] In this invention, the server includes means for acquiring information in real time from monitoring devices installed on-site, means for processing the acquired information and detecting unsafe behavior or conditions using a generation algorithm, and means for evaluating the risk based on the detected unsafe behavior or conditions according to set criteria and notifying the administrator. This enables real-time detection of unsafe conditions and rapid notification of risks.
[0600] A "monitoring device" is a device installed on-site that acquires real-time information on workers' actions and the environment.
[0601] A "generative algorithm" is a computational method that learns from past data and automatically identifies unsafe behaviors and states.
[0602] A "manager" is a person responsible for on-site safety management, receiving risk information, and taking appropriate action.
[0603] "Hazard" refers to the potential risks that arise when work or the work environment at a site deviates from safety standards.
[0604] "Notification" is the procedure for informing administrators about detected unsafe situations and the countermeasures being taken.
[0605] As an embodiment of the present invention, a method for constructing a system to enhance safety management will be described. This system provides comprehensive safety management with monitoring devices, servers, and terminals as its main components.
[0606] Hardware and software usage
[0607] The monitoring system is a device that combines cameras and sensors to monitor the movements of workers and the surrounding environment in real time. The information obtained from these devices is transmitted to a server.
[0608] The server functions as the central processing unit. This server is equipped with a generative AI model, which has been trained on data collected under various historical circumstances. This model analyzes data from the field and identifies unsafe behaviors and conditions.
[0609] Information about detected dangerous conditions is transmitted from the server to the terminals. These terminals consist of mobile devices and computers, allowing administrators to monitor the situation on-site and plan countermeasures in real time.
[0610] Specific examples and prompt statements
[0611] For example, at a construction site, if a monitoring device detects that a worker is not wearing the necessary safety equipment while working at height, the server uses a generated AI model to recognize this as an unsafe situation. The terminal notifies the administrator that "safety equipment is not being worn while working at height," prompting immediate action. In this case, the terminal will also provide an alert sound or vibration as needed.
[0612] Examples of prompts used to make a generative AI model work include the following:
[0613] "Monitor the work site in real time and determine whether safety standards are being met."
[0614] "Detect unsafe behaviors on-site and assess the risk level."
[0615] As described above, the system of the present invention surpasses conventional methods, enabling faster and more effective safety management.
[0616] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0617] Step 1:
[0618] The server acquires information in real time from monitoring devices installed on-site. These monitoring devices use cameras and sensors to capture worker movements and surrounding environment information, and transmit this data to the server. This data is used as input for analysis by the generated AI model.
[0619] Step 2:
[0620] The server analyzes the received information using a generating AI model. This model is trained on historical data and executes algorithms to identify unsafe behaviors and states. Specifically, it performs image analysis and behavioral pattern recognition, and compares the results with a database. As a result, data identified as unsafe is generated as output.
[0621] Step 3:
[0622] Unsafe conditions detected by the server are evaluated by a risk assessment algorithm. The server quantifies the risk against established criteria, and the evaluation results become the input for the next step. This reveals that high-priority situations require a rapid response.
[0623] Step 4:
[0624] Based on the evaluation results, the server sends a notification to the administrator via the terminal. The notification clearly states the details of the detected condition and the need for immediate action depending on the urgency. Specifically, it uses alert sounds and vibration functions to draw the administrator's attention and displays video information on the terminal.
[0625] Step 5:
[0626] Based on notifications and countermeasures received from the terminal, the user provides guidance to on-site workers. The server displays specific improvement measures and training information on the terminal, and the user uses this information to provide feedback to the workers. This process aims to improve safety awareness. In addition, if there are similar past cases, information referenced from the database is also added.
[0627] (Application Example 1)
[0628] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0629] Safety management at the workplace is extremely important for protecting the lives and health of workers. However, conventional safety management systems have problems such as not being able to quickly detect and warn of unsafe conditions in real time, and workers having difficulty obtaining necessary safety information immediately. Therefore, there is a need for a safety management system that can effectively provide immediate warnings and education to workers.
[0630] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0631] In this invention, the server includes means for acquiring video information in real time from observation devices installed on site, means for processing the acquired video information and detecting unsafe actions or conditions using a generation model, and means for presenting the safety status to workers in real time using a visual device and issuing audio and visual warnings as necessary. This enables immediate assessment of safety conditions at the work site and rapid warning.
[0632] An "observation device" is a piece of equipment installed on-site to acquire video information in real time.
[0633] "Visual information" refers to visual data acquired from observation equipment via cameras and sensors.
[0634] A "generative model" is a machine learning algorithm used to process video information and recognize and detect unsafe actions or states.
[0635] "Risk assessment" is the process of determining the severity and urgency of a risk based on detected unsafe behaviors or conditions, using predetermined criteria.
[0636] A "safety management officer" is responsible for safety management at the work site, receiving risk assessment results, and taking appropriate action.
[0637] A "visual device" is a device worn by workers that is used to visually present video information or warnings.
[0638] "Audio and visual warnings" are notification methods using sound and images that are used to inform workers of unsafe conditions.
[0639] The system implementing the present invention comprises an observation device, a server, and a visual device to enhance safety management at the work site. The observation device is installed at the site and monitors the worker's movements and the work environment in real time using cameras and sensors. Real-time video information transmitted from the observation device is received by the server, which uses a generative model based on this video information to detect unsafe movements and conditions.
[0640] The server analyzes video information in detail and can utilize TensorFlow or PyTorch as its generative model. This allows it to automatically identify various unsafe conditions. If the server detects an unsafe condition, it performs a risk assessment based on that information and determines the severity of the risk according to predetermined criteria.
[0641] Next, the server immediately notifies the worker of the risk assessment results through a visual device. Smart glasses are used as the visual device, and by issuing audio and visual warnings, the worker can immediately understand the hazard and take the necessary measures. Specific examples of devices that serve as this interface include Google Glass and Microsoft HoloLens.
[0642] As a concrete example, consider a scenario where a worker accidentally enters an area outside of the designated control zone at a work site. In this case, the monitoring device detects the unsafe behavior, the server analyzes it, and immediately sends a warning to the smart glasses. This warning prompts the worker to quickly return to the safe area.
[0643] An example of a prompt might be: "Input frame images of the manufacturing area into the AI model and evaluate safety. If there are obstacles in the worker's path, issue a warning to the user. Provide suggestions for safety measures before proceeding to the next frame." This prompt provides important guidance for the AI model to support safety management in the workplace.
[0644] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0645] Step 1:
[0646] The observation device acquires video information from the work site. This device uses a camera to capture the entire work environment in real time and transmits the acquired video as digital data to a server. The input is raw video from the camera, and the output is digitized image data.
[0647] Step 2:
[0648] The server receives video data from the observation device. The server divides the received image data into frames and inputs them into a generative AI model. In this process, the video data is compared with a database and features are extracted. The input is the divided image frames, and the output is a feature vector.
[0649] Step 3:
[0650] The server uses a generative AI model to analyze the input feature vector and detect unsafe behaviors and states. Based on the training data, the generative AI model identifies potential risks in the video. The input is a feature vector, and the output is the risk assessment result.
[0651] Step 4:
[0652] The server performs a risk assessment on detected risks based on established criteria and classifies the results in terms of importance and urgency. This assessment calculates the severity of the risk, and the appropriate action is determined according to the urgency. The input is the risk assessment result, and the output is the risk assessment score.
[0653] Step 5:
[0654] The server is connected to a visual device and communicates the assessed risk to the worker. Based on the risk score, the visual device presents the worker with audio and visual alerts. This process allows the worker to immediately take appropriate safety actions. The input is the risk assessment score, and the output is the alert notification.
[0655] Step 6:
[0656] The user (safety manager) reviews the information received via the visual device and issues instructions to the worker. Specifically, they instruct the worker to implement appropriate safety measures and provide ongoing safety training. The input is an alert notification, and the output is instructions and training information for the worker.
[0657] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0658] This invention is a system that combines observation devices, a server, a terminal, and an emotion engine to enhance on-site safety management and raise safety awareness among workers. This system is characterized not only by its ability to detect unsafe actions and conditions in real time, but also by its ability to recognize user emotions and respond in a more personalized manner.
[0659] The system consists of observation devices installed on-site to capture video footage of workers' movements and surrounding conditions. A server receives this video data and uses a generative model to detect unsafe actions or conditions. At this time, a risk assessment is performed, and a notification based on the assessment results is sent to the terminal.
[0660] Furthermore, the addition of an emotion engine allows for the acquisition of information about the user's emotional state. The server analyzes this emotional data through the camera and microphone on the user's device and incorporates it into risk assessments and notification content. For example, if a user is experiencing high levels of stress, notifications can be made simpler and more specific to improve the efficiency of the response.
[0661] The device displays notifications and suggestions received from the server to the user. Because the presentation method is adjusted based on sentiment data, users receive information optimized according to their own emotional state.
[0662] For example, if a user is a safety manager at a work site and discovers a worker not wearing a safety harness while working at height, the server detects this risk and uses an emotion engine to determine that the user is in a state of anxiety. The terminal then displays countermeasures and past examples in a fast and easy-to-understand format. This process allows the user to give instructions quickly and calmly, which is expected to improve safety at the work site.
[0663] The following describes the processing flow.
[0664] Step 1:
[0665] The server receives real-time video data from observation devices installed on-site. This data includes worker movements and surrounding conditions and is used for safety assessments.
[0666] Step 2:
[0667] The server preprocesses the received video data and uses a generative model to detect unsafe behaviors and states. This model allows for the automatic identification of behaviors deemed dangerous.
[0668] Step 3:
[0669] When the server detects unsafe behavior or conditions, it performs a risk assessment. The assessment is based on pre-defined criteria and similar past cases, and determines the risk level of the detected event.
[0670] Step 4:
[0671] The server acquires emotional data from the device to recognize the user's emotions. It uses the device's camera and microphone to analyze the user's emotional state from their voice tone and facial expressions.
[0672] Step 5:
[0673] The server generates notifications and suggestions based on risk assessment results and the user's emotional state. When emotions are heightened, it strives to provide information that is more concise and easy to understand.
[0674] Step 6:
[0675] The device receives notifications sent from the server and displays them on the screen. These notifications include countermeasures against unsafe behavior and risk information that is sensitive to the user's feelings.
[0676] Step 7:
[0677] Users review notifications displayed on their devices and provide appropriate instructions to field workers based on the situation. Users can make calm and quick decisions based on information optimized for their emotions.
[0678] (Example 2)
[0679] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0680] On-site safety management requires the detection and prompt response to unsafe actions and conditions. However, conventional technologies only detect and notify of hazards, failing to consider the emotional state of workers and managers, and sometimes resulting in inaccurate and swift responses. Furthermore, there is a challenge in effectively referencing past cases of dangerous situations to support rapid decision-making on-site.
[0681] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0682] In this invention, the server includes means for acquiring image information in real time from a detection device installed on-site, means for processing the acquired image information and detecting dangerous actions or states using a generation artificial intelligence model, and means for analyzing the user's emotional information and adjusting the risk assessment and notification content based on the user's emotional state. This enables more rapid and appropriate management of dangerous situations in conjunction with emotional information, and allows for the provision of educational information that takes past cases into consideration.
[0683] A "detection device" is a device installed on-site to acquire image information in real time.
[0684] "Image information" refers to visual data that shows the actions of workers and the conditions of the environment.
[0685] A "generative artificial intelligence model" is a machine learning model trained to detect dangerous actions or conditions using acquired image information.
[0686] A "hazardous action or condition" is any action or surrounding condition that could pose a potential risk to workers or equipment.
[0687] "Risk assessment" is the process of estimating the degree of risk based on detected hazardous actions or conditions.
[0688] "User emotional information" refers to data that indicates the user's emotional state, and is obtained from information such as voice tone and facial expressions.
[0689] "Educational information" refers to instructional data provided to workers, including suggestions for improvement regarding dangerous actions or conditions.
[0690] A "safety management officer" is a person responsible for maintaining safety at the worksite and taking appropriate measures.
[0691] "Example information" refers to case study and example data related to specific dangerous actions or conditions, extracted from past information resources.
[0692] This invention is a system that utilizes a server, terminals, detection devices, and a generative AI model to enhance on-site safety management. The server processes real-time image information acquired from on-site detection devices and uses the generative AI model to detect dangerous actions or conditions. This system incorporates an emotion analysis engine that analyzes the user's emotional information and reflects it in notification content and risk assessment.
[0693] Specifically, the server utilizes hardware that receives data from cameras and sensors installed on-site. The received data is analyzed using generative artificial intelligence models that detect dangerous behaviors and states, employing machine learning frameworks such as TensorFlow and PyTorch. In addition, the sentiment analysis engine analyzes the user's emotional information using data collected from the user's device via camera and microphone.
[0694] If the user is a safety manager at the work site, the system will send a notification to the terminal that takes the user's emotional state into consideration when a specific hazard is detected. For example, if the hazard of not wearing safety equipment during work at height is detected and the user is determined to be in a state of agitation, a quick and clear suggestion will be displayed on the terminal. An example of such a prompt message would be: "Unsafe behavior observed at a work site at height. The user is in a state of agitation. Please briefly outline risk mitigation measures."
[0695] This allows users to respond quickly and accurately to hazards, thereby improving safety at the worksite.
[0696] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0697] Step 1:
[0698] The server receives image information in real time from on-site detection devices. This image information includes worker movements and environmental conditions. The received image data is preprocessed by noise reduction and conversion to an appropriate resolution, and then processed into a format suitable for analysis by AI models. The output is the processed image data.
[0699] Step 2:
[0700] The server passes processed image data as input to a generating artificial intelligence model to detect dangerous behaviors and states. This model is a deep learning model built using, for example, TensorFlow, and has already learned about dangerous behaviors and states. Data calculations include feature extraction and classification, and the degree of danger is evaluated. The output is a list of dangerous behaviors and states detected.
[0701] Step 3:
[0702] The server uses the camera and microphone on the user's device to collect emotional information in order to analyze the user's emotional state based on the risk detection results. This data includes facial recognition and voice tone analysis. For data processing, emotional characteristics are extracted from the video and audio and evaluated using an emotion analysis model. The output is the user's emotional state data.
[0703] Step 4:
[0704] The server integrates hazard detection results and emotional state data to generate notification content. This generation process creates prompt messages based on the urgency of the hazard and the user's emotional state. For example, it might say, "Unsafe behavior has been detected during work at height, and the user is in a state of anxiety. A concise immediate risk mitigation plan is presented." The output is a customized notification message.
[0705] Step 5:
[0706] The terminal displays notification messages sent from the server to the user. Here, specific actions are taken to adjust how notifications are displayed based on the user's emotional state. For example, for a user in an anxious state, larger fonts and high-contrast colors are used to make important information stand out. The output is an optimized information presentation for the user.
[0707] (Application Example 2)
[0708] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0709] Accidents caused by worker negligence or errors are frequent at work sites, leading to decreased productivity and personal injury. In this situation, there is a need for more efficient and accurate methods to ensure worker safety and raise safety awareness at the workplace. Furthermore, the lack of appropriate responses to workers' emotional states can hinder calm judgment.
[0710] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0711] In this invention, the server includes means for acquiring time-series data from observation devices installed in the work environment; means for processing the acquired time-series data and identifying dangerous behaviors or situations using a generative model; means for evaluating the danger based on defined criteria and notifying information providers based on the identified dangerous behaviors or situations; and means for collecting worker emotional information using an emotion analysis mechanism and adjusting the notification content based on the evaluation results. This enhances worker safety while enabling rapid on-site response.
[0712] "Observation devices installed in the work environment" are devices placed at the work site to grasp the conditions and activities of workers in that location in real time and collect them as time-series data.
[0713] "Time-series data" refers to data collected over time, allowing for sequential monitoring of the work environment and the status of workers.
[0714] A "generative model" is an algorithm that is trained on large amounts of data and used to identify specific patterns or abnormal behavior.
[0715] "Dangerous behavior or situation" refers to any action or condition within the work environment that increases the risk to workers or equipment, and includes those requiring safety management precautions.
[0716] "Defined criteria" are pre-established rules or conditions for assessing risks and determining appropriate responses.
[0717] An "information provider" is an individual or organization that is involved in safety management and is responsible for ensuring the safety of workers and equipment.
[0718] An "emotional analysis mechanism" is a technology or device used to analyze the emotional state of workers and acquire it as data.
[0719] "Adjusting notification content" refers to the process of changing or suppressing warning or guidance information depending on the emotional state of the worker.
[0720] The system that realizes this invention consists of an observation device, a server, a terminal, and an emotion analysis mechanism. First, the observation device is installed in the work environment and acquires time-series data of the worker and the work site. This data is transferred to the server via the network.
[0721] The server preprocesses data using image processing libraries such as OpenCV and identifies dangerous behaviors or situations by running generative AI models utilizing machine learning libraries such as TensorFlow. It also uses software such as Microsoft Azure Cognitive Services to analyze video and audio data received from terminals and determine the emotional state of the worker. Based on the identified dangerous behaviors, the server evaluates the risk level and sends a customized notification to the terminal of the information provider (the worker or manager who is the user).
[0722] The terminal displays received notifications appropriately to the worker. The displayed content is adjusted according to the worker's emotional state; for example, a message encouraging calmness is added to a worker who is feeling anxious. In addition, support is provided to help workers make better decisions by referring to past case information.
[0723] As a concrete example, imagine a worker performing tasks next to a conveyor belt in a factory. The server performs emotion analysis, and if it detects that the worker is stressed, it displays a message on the smart glasses saying, "Assess the situation and proceed calmly to the next step."
[0724] Examples of prompts for the generating AI model include: "Propose a method to detect unsafe behavior in real time, analyze user sentiment data, and immediately display suggested responses." This prompt allows the system to generate appropriate information to support worker safety and efficient work.
[0725] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0726] Step 1:
[0727] The server receives time-series data from observation equipment installed in the work environment. This data includes camera images and sensor information. The server preprocesses the video data using OpenCV to remove noise and adjust the resolution. This allows for the acquisition of clean image data suitable for analysis.
[0728] Step 2:
[0729] The server analyzes the pre-processed video data using TensorFlow. It then runs a generative AI model to identify dangerous behaviors or situations in real time. In this step, the video data is given as input to the model, and the detected dangerous situations, their characteristics, and risk levels are returned as output.
[0730] Step 3:
[0731] The server uses Microsoft Azure Cognitive Services to perform emotion analysis. It takes audio and video data received from the worker's terminal as input and analyzes the worker's emotional state. The analysis results are output as an emotional state assessment (e.g., tension, anxiety, relaxation).
[0732] Step 4:
[0733] The server integrates identified risky behaviors and sentiment analysis results to create notifications for informants. These notifications are customized based on the risk assessment and emotional state of the generating AI model. For example, the notification might include a specific message such as, "Please calmly assess the situation and proceed cautiously with the next steps."
[0734] Step 5:
[0735] The terminal displays notifications received from the server to the user (worker). This display is done via smart glasses or a smartphone. By receiving these notifications, workers can obtain guidance for taking appropriate action.
[0736] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0737] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0738] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0739] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0740] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0741] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0742] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0743] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0744] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0745] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0746] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0747] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0748] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0749] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0750] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0751] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0752] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0753] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0754] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0755] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0756] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0757] The following is further disclosed regarding the embodiments described above.
[0758] (Claim 1)
[0759] A means of acquiring video information in real time from observation equipment installed on site,
[0760] A means for processing acquired video information and detecting unsafe behavior or states using a generative model,
[0761] A means of evaluating the risk based on the detected unsafe operation or condition according to predetermined criteria and notifying the safety management officer,
[0762] A system that includes this.
[0763] (Claim 2)
[0764] The system according to claim 1, which generates improvement suggestions for detected unsafe operations or conditions and provides them to workers as educational information.
[0765] (Claim 3)
[0766] The system according to claim 1, which refers to a past database and extracts case information related to detected unsafe behavior or conditions.
[0767] "Example 1"
[0768] (Claim 1)
[0769] A means of acquiring information in real time from monitoring devices installed on site,
[0770] A means for processing acquired information and detecting unsafe behavior or states using a generation algorithm,
[0771] A means of assessing the risk based on detected unsafe behavior or conditions according to established criteria and notifying the administrator,
[0772] Based on the information received, specific countermeasures will be presented and provided to the person in charge as educational information via a terminal.
[0773] A system that includes this.
[0774] (Claim 2)
[0775] The system according to claim 1, which refers to a collection of past information and extracts case information related to detected unsafe behavior or conditions.
[0776] (Claim 3)
[0777] The system according to claim 1, which uses video, audio, or vibration to convey the urgency of a notification.
[0778] "Application Example 1"
[0779] (Claim 1)
[0780] A means of acquiring video information in real time from observation equipment installed on site,
[0781] A means for processing acquired video information and detecting unsafe behavior or states using a generative model,
[0782] A means of evaluating the risk based on the detected unsafe operation or condition according to predetermined criteria and notifying the safety management officer,
[0783] A means of presenting safety conditions to workers in real time using visual devices and issuing audio and visual warnings as needed,
[0784] A system that includes this.
[0785] (Claim 2)
[0786] The system according to claim 1, which generates improvement suggestions for detected unsafe operations or conditions and provides them to workers as educational information.
[0787] (Claim 3)
[0788] The system according to claim 1, which refers to a past database and extracts case information related to detected unsafe behavior or conditions.
[0789] "Example 2 of combining an emotion engine"
[0790] (Claim 1)
[0791] A means of acquiring image information in real time from a detection device installed on site,
[0792] A means for processing acquired image information and detecting dangerous actions or states using an artificial intelligence model,
[0793] A means of evaluating the degree of risk based on detected dangerous actions or conditions according to predetermined criteria and notifying the safety management officer,
[0794] A means for analyzing user emotional information and adjusting risk assessment and notification content based on the user's emotional state,
[0795] A system that includes this.
[0796] (Claim 2)
[0797] The system according to claim 1, which generates appropriate improvement suggestions based on detected dangerous actions or conditions and the user's emotional state, and provides them to the worker as educational information.
[0798] (Claim 3)
[0799] The system according to claim 1, which refers to past information resources, extracts illustrative information related to detected dangerous behavior or conditions, and provides illustrative information according to the user's emotional state.
[0800] "Application example 2 when combining with an emotional engine"
[0801] (Claim 1)
[0802] A means of acquiring time-series data from observation equipment installed in the work environment,
[0803] A means of processing acquired time-series data and using a generative model to identify dangerous behaviors or situations,
[0804] A means of assessing the risk based on identified dangerous behaviors or situations according to defined criteria and notifying the informant,
[0805] A means of collecting workers' emotional information using an emotion analysis mechanism and adjusting notification content based on the evaluation results,
[0806] A system that includes this.
[0807] (Claim 2)
[0808] The system according to claim 1, which creates improvement plans for identified dangerous behaviors or situations and provides them to workers as educational resources.
[0809] (Claim 3)
[0810] The system according to claim 1, which refers to a collection of case examples and extracts case information related to identified dangerous behaviors or situations. [Explanation of Symbols]
[0811] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring video information in real time from observation equipment installed on site, A means for processing acquired video information and detecting unsafe behavior or states using a generative model, A means of evaluating the risk based on the detected unsafe operation or condition according to predetermined criteria and notifying the safety management officer, A system that includes this.
2. The system according to claim 1, which generates improvement suggestions for detected unsafe operations or conditions and provides them to workers as educational information.
3. The system according to claim 1, which refers to a past database and extracts case information related to detected unsafe behavior or conditions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A