system
The system addresses delays in disaster response by using terminal devices to collect and analyze acoustic and image signals, enabling rapid and accurate rescue instruction generation for efficient disaster management.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing disaster response systems face delays in information collection and labor-intensive judgment processes, lacking rapidity and accuracy in generating rescue instructions, which hampers efficient rescue activities.
A system that utilizes terminal devices to collect on-site acoustic and image signals, transmitting them to a central processing unit for analysis, which generates and communicates rescue instructions efficiently and accurately using generative AI and machine learning algorithms.
Enables faster and more efficient disaster response by providing real-time, accurate rescue instructions, reducing damage and improving safety.
Smart Images

Figure 2026073415000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a need for an efficient system for quickly and accurately collecting information from the disaster site during a disaster and generating and communicating rescue instructions based on that information. Conventional methods have problems such as delays in information collection and a large amount of labor for judgment, lacking the rapidity and accuracy of rescue activities. There is a need for technology to solve this problem and perform rescue activities efficiently.
Means for Solving the Problems
[0005] This invention provides a system that collects on-site environmental data using terminal devices that acquire acoustic and image signals, and transmits the collected data to a central processing unit via a communication network. The central processing unit analyzes the received signals and extracts important information. Based on this information, it generates optimal rescue instructions and transmits them to the terminal devices, which then display the instructions to the user, providing a system that supports rescue activities. This invention aims to expedite and improve the accuracy of information gathering and rescue instructions during disasters, thereby enhancing the efficiency of rescue activities.
[0006] An "acoustic signal" is a signal obtained by converting sound information transmitted through air or other media into an electrical form.
[0007] An "image signal" is a digital or analog signal that electronically represents visual information.
[0008] A "terminal device" is a device that has the function of acquiring acoustic and image signals and transmitting data via a communication network.
[0009] A "central processing unit" is a computer system that analyzes received acoustic and image signals and extracts important information.
[0010] A "communication network" refers to the entire network infrastructure and lines used to transmit data between multiple locations.
[0011] "Generative artificial intelligence" is a system that uses machine learning and algorithms to automatically perform complex tasks and make decisions.
[0012] A "machine learning algorithm" is a computational method that learns patterns based on data and performs inference and prediction.
[0013] A "relief order" is a command that outlines the necessary guidelines for carrying out efficient relief activities during a disaster. [Brief explanation of the drawing]
[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is a system that collects on-site environmental data using terminal devices that acquire acoustic and image signals, and transmits this data to a central processing unit via a communication network. The system includes the following program:
[0036] The terminal acquires acoustic and image signals and transmits this data to a server in real time. For example, a terminal carried by a rescue worker at a disaster site captures the surroundings and sounds through a camera and microphone mounted on their helmet. This allows the situation at the site to be acquired as digital data in real time.
[0037] The server analyzes the received acoustic and image signals. Using generative AI and machine learning algorithms, the server processes large amounts of data quickly and accurately, identifying damage and locations with a high probability of human casualties. This analysis allows for the extraction of information necessary to determine rescue priorities.
[0038] The server generates rescue orders from the extracted information, including guidelines on what kind of rescue operations should be carried out and where. For example, if it discovers a collapsed building from video data of a disaster area and, after analyzing the evacuation situation in the surrounding area, determines that there is a high probability that residents have not yet evacuated, it will issue an order for rapid rescue operations to that location. In this way, the server always generates accurate rescue orders based on the latest information.
[0039] Users receive these rescue instructions through their devices and carry out rescue operations based on them. The devices display the instructions sent from the server on the screen and notify the user via voice. This allows rescue workers to act efficiently and effectively. For example, users can check evacuation points and hazardous areas displayed on a map and act according to the instructions.
[0040] This system will enable faster and more efficient relief efforts during disasters, which is expected to reduce damage and improve safety.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] The device acquires audio and image signals.
[0044] The device is installed on-site, capturing surrounding video with its camera and recording ambient sounds with its microphone. This data is temporarily stored within the device.
[0045] Step 2:
[0046] The device sends data to the server.
[0047] The terminal transmits the acquired acoustic and image signals to the server in real time via the communication line. The data is encrypted to ensure the security of the transmission.
[0048] Step 3:
[0049] The server receives the data and begins analysis.
[0050] The server stores the received data in an analysis database and uses a generating AI to analyze the data. In this process, images are used to identify collapsed buildings and the presence or absence of survivors, and voices and rescue requests are extracted from audio.
[0051] Step 4:
[0052] The server extracts important information and generates rescue orders.
[0053] The server uses the analyzed data to assess the damage and evacuation situation in specific areas. Based on this information, it determines what kind of relief is needed in each area and creates specific relief orders.
[0054] Step 5:
[0055] The server sends a rescue order to the terminal.
[0056] The server sends the generated rescue instructions to the relevant terminals. This allows users receiving the instructions on-site to take immediate action.
[0057] Step 6:
[0058] The user checks and executes the rescue order on their device.
[0059] The user checks the instructions notified on their device and understands the details of the instructions from the on-screen interface. Then, they quickly carry out rescue operations based on the instructions.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] Disaster response activities on the ground require swift and accurate decision-making, but there is a lack of technology to collect reliable information in real time and provide timely instructions. Furthermore, relying on manual analysis of the collected information can lead to delays in response. This challenge poses a significant obstacle to minimizing damage and enhancing the effectiveness of relief efforts.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes means for collecting on-site environmental information using an information terminal that acquires sound and images, means for transmitting the acquired sound and images to a central processing unit via a communication path, and means for analyzing the received information in the central processing unit and extracting important content. This makes it possible to quickly calculate the priority of rescue efforts and generate and transmit optimal rescue guidelines.
[0065] "Acoustics" refers to signals based on sound waves generated from the environment and objects, and serves as a medium for obtaining information through these signals.
[0066] An "image" is digital data that represents visual information and is collected through cameras or other visual sensors.
[0067] An "information terminal" is a computer or electronic device used to acquire data such as sound and images and transmit it to other devices via a communication network.
[0068] A "communication path" is a path that includes the infrastructure and network technologies used to send and receive data.
[0069] A "central processing unit" is a computer system or server used to process and analyze received data.
[0070] "Generative artificial intelligence" is a technology or algorithm that uses machine learning based on large amounts of data to enable pattern recognition and automated analysis.
[0071] "Machine learning methods" are models and algorithms that automatically learn from data and perform inference and decision-making.
[0072] "Relief guidelines" are a set of action plans and instructions given to ensure that relief operations are carried out efficiently and effectively.
[0073] The system according to this invention consists mainly of three components: a terminal, a server, and a user.
[0074] First, the terminal is responsible for acquiring audio and video signals on-site. Specifically, a device equipped with a camera and microphone is used. This terminal continuously captures sound and video, and digitizes them in real time. This data is transmitted to a server via a communication network.
[0075] Next, the server analyzes the received data. The server possesses advanced computing power and uses generative AI models and machine learning algorithms to analyze the received acoustic and image signals in detail. This analysis process helps to understand the situation on site and identify high-risk areas and points requiring rescue. For example, the server can identify whether buildings have collapsed or to distinguish between different patterns of cries for help.
[0076] Finally, a key feature of this system is that it provides users with rescue guidelines based on the analysis results. The terminal receives instructions from the server, displays them to the user, and provides voice notifications as needed. The user can then quickly carry out rescue activities based on the information received. For example, they can move according to evacuation guidelines displayed by a map application.
[0077] This system provides rapid and accurate information processing in disaster response. A specific example of its use is inputting a prompt message to the server such as, "Analyze the latest acoustic and image data transmitted from the disaster area and identify locations where human casualties are likely." Based on this prompt message, the server performs the necessary analysis and generates results indicating the optimal relief strategy.
[0078] This invention provides a system that enables more efficient and rapid disaster response activities, thereby minimizing damage.
[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0080] Step 1:
[0081] The terminal acquires acoustic and image signals from the site. It captures ambient sound and video as input and converts them into digital data. As output, it generates real-time aggregated digital acoustic and image data.
[0082] Step 2:
[0083] The terminal transmits the generated audio and image data over the communication network. In this step, it performs operations to properly packetize the digital data and send it to the server. The input is the data generated in step 1, and the output is the data transmission to the server.
[0084] Step 3:
[0085] The server analyzes the acoustic and image data received from the terminal. Using generative AI models and machine learning algorithms, it extracts important features from the input data. Specifically, it processes data to identify building collapses, human voices, and unusual sound patterns. As output, it generates information identifying high-risk areas and points requiring rescue.
[0086] Step 4:
[0087] The server generates rescue guidelines based on the information extracted in step 3 and sends them to the terminal. Based on the results of the generated AI, it forms action guidelines for high-priority rescue activities. Using the identified risk information as input, it creates specific rescue guidelines as output and sends them to the terminal.
[0088] Step 5:
[0089] Users initiate action using rescue guidelines received through their devices. The devices display the guidelines and provide audio alerts as needed. Input is rescue guidelines sent from the server, and output is information provided to the user through screen displays and audio notifications. Based on the information displayed on the devices, users quickly carry out evacuation and support activities.
[0090] (Application Example 1)
[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0092] In modern society, large commercial facilities and event venues attract large crowds, potentially creating dangerous situations. Conventional monitoring systems struggle with real-time anomaly detection and rapid response, necessitating more efficient and effective solutions to these challenges.
[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0094] In this invention, the server includes means for collecting on-site environmental information using an information acquisition device that acquires acoustic and image signals, means for transmitting the acquired acoustic and image signals to a data processing device via a communication network, and means for analyzing the received signals and extracting important information, including abnormal events, in the data processing device. This enables real-time anomaly detection and rapid response in commercial facilities and event venues.
[0095] An "information acquisition device" is a device that acquires acoustic and image signals on-site and is used to collect environmental information.
[0096] A "communication network" is a communication system used to transmit signals acquired from an information acquisition device to a data processing device.
[0097] A "data processing device" is a device that analyzes received acoustic and image signals and extracts important information.
[0098] An "abnormal event" is information indicating a situation that deviates from the normal state or operation, and is an event that requires particular attention in a monitoring system.
[0099] "Important information" refers to data, including abnormal events, that is essential for real-time decision-making and response.
[0100] "Action instructions" are instructions generated based on extracted key information, and are displayed to the user to prompt specific actions.
[0101] In the system of the present invention, "user" refers to a person or organization that receives action instructions transmitted to the information acquisition device and actually takes action.
[0102] The system of the present invention provides a secure environment in commercial facilities and event venues by combining an information acquisition device, a communication network, and a data processing device. The information acquisition device acquires acoustic and image signals in real time using a camera and microphone mounted on smart glasses or a smartphone. These signals are transmitted to the data processing device via the communication network.
[0103] The data processing unit uses generative artificial intelligence models such as TENSORFLOW® and PyTorch, built using Python, to analyze received acoustic and image signals. This analysis allows for the rapid extraction of abnormal events and important information from the environment. For example, data processing is performed to detect abnormal behavioral patterns or changes in volume.
[0104] The server generates appropriate action instructions based on these abnormal events and transmits them to the information acquisition device. The information acquisition device displays the action instructions to the user and also provides notification via audio output. This enables the user to understand the situation in real time and respond quickly.
[0105] As a concrete example, this system could be used at large-scale events. Multiple information acquisition devices could be installed at the entrance of the event venue and in crowded areas to monitor the movement of attendees in real time. In this case, by using a generative AI model to set example prompts such as "If loud voices are heard from attendees, it will be judged as an anomaly and security guards will be notified," it would be possible to immediately instruct countermeasures in the event of an anomaly.
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The terminal captures acoustic and image signals in real time using the camera and microphone of smart glasses or a smartphone to acquire on-site environmental information. The input is raw data from the site, and the output is digital signal data.
[0109] Step 2:
[0110] The terminal transmits the acquired acoustic and image signals as digital data to the data processing device via a communication network. The input is digitized acoustic and image data, and the output is communication data to the data processing device.
[0111] Step 3:
[0112] The server, in its data processing unit, runs a generative artificial intelligence model using TensorFlow or PyTorch to analyze received signals. The input is the transmitted signal data, and the output is anomalous events and important information extracted through the analysis. This analysis can, for example, detect peaks in audio waveforms or evaluate the degree of crowding from images.
[0113] Step 4:
[0114] The server generates appropriate action instructions based on the abnormal events obtained through analysis. The input is information about the abnormal events, and the output is specific action instructions. Using a generation AI model, instructions are generated according to prompt statements such as "If congestion is high in some areas, instruct security guards to check those areas."
[0115] Step 5:
[0116] The server transmits the generated action instructions to the information acquisition device, and the terminal notifies the user of these instructions. The input is action instruction data from the server, and the output is visual and auditory notifications to the user. The terminal displays the instructions on its screen and issues an audible warning.
[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0118] The system in this invention not only collects on-site environmental data using terminal devices that acquire acoustic and image signals, but also incorporates an emotion engine that recognizes the user's emotions. This system is designed to enable effective rescue operations at disaster sites.
[0119] The device captures the voice and facial expressions of users in the disaster area and sends this data to a server. This allows the device to monitor the user's emotional state in real time and obtain emotion-based information as needed. For example, if a member of a rescue team is experiencing high stress levels, this information is recorded as emotional data.
[0120] The server analyzes received emotional data and environmental data acquired by sensors. By utilizing generative artificial intelligence and machine learning algorithms, the server automatically recognizes the situation from this data and determines the appropriate course of action. Emotional data is used to re-evaluate the priority of rescue efforts, particularly in cases of high stress levels or urgency.
[0121] Based on these analysis results, the server generates and provides the user with the most appropriate rescue instructions. For example, if the emotion engine highly values the user's anxiety, the server sends information about psychological support and urgent encouraging messages to the device. The user receives this immediately and can use it to guide their next actions.
[0122] The terminal displays rescue instructions and emotion-based support information sent from the server to the user, and also provides voice guidance. Users can review the instructions and confidently take appropriate action. This system not only streamlines rescue operations but also contributes to maintaining the user's mental well-being.
[0123] Thus, the system of the present invention, which combines an emotion engine, is an important solution for supporting disaster relief activities in a multifaceted way and maintaining a safe and secure environment.
[0124] The following describes the processing flow.
[0125] Step 1:
[0126] The device acquires audio and image signals.
[0127] The terminals are deployed on-site to capture the user's voice and facial expressions in real time. Using cameras and microphones, they collect surrounding video and audio data, while simultaneously acquiring emotion-related data such as voice tone and changes in facial expressions.
[0128] Step 2:
[0129] The device analyzes the data using an emotion engine and sends the results to the server.
[0130] The emotion engine within the device determines the user's emotional state from collected acoustic and image signals. For example, it assesses stress levels from voice tone and speaking style, and detects signs of sadness or anxiety through facial expression analysis. The data, including these analysis results, is then transmitted to a server via the communication network.
[0131] Step 3:
[0132] The server analyzes all the data and extracts important information.
[0133] The server integrates and analyzes emotional and environmental data. Generative AI assesses the specific situation in the disaster area and determines which areas should be prioritized. By considering emotional data, it identifies the location and circumstances of users who particularly need psychological support.
[0134] Step 4:
[0135] The server generates optimal rescue instructions and psychological support information.
[0136] Based on the analysis results, the server determines the necessary relief measures and mental health support information. For example, the server might create instructions for a high-stress user, including resources for psychological counseling and encouraging messages.
[0137] Step 5:
[0138] The server generates instructions and sends them to the terminal.
[0139] The server sends rescue instructions and psychological support information to the terminal, enabling users to respond quickly. The data is encrypted and received instantly.
[0140] Step 6:
[0141] Users receive information and take action through their devices.
[0142] Users review the rescue instructions and support messages displayed on the device. This allows users to quickly take appropriate rescue actions while managing their emotional state. The device also provides voice guidance to help understand the instructions.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] In disaster areas, it is essential to grasp various emergency information and the emotional state of users in real time and to appropriately direct rescue operations. However, existing systems have problems in that they cannot adequately process information accurately and respond quickly while taking into account the emotional information of users. Therefore, it is necessary to develop a system that centrally processes diverse data, including emotional information, and provides more appropriate support instructions based on that information.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes means for collecting on-site situation information using an information terminal that acquires auditory and visual data, means for transmitting the acquired auditory and visual data to a centralized information processing device via a data transmission path, and means for analyzing the received data and identifying and extracting important elements in the information processing device. This enables rapid and accurate understanding of various situations at disaster sites and allows for appropriate support instructions that take into account the emotional state of users.
[0148] "Auditory data" refers to information that includes acoustic signals, such as speech and ambient sounds, acquired in digital format.
[0149] "Visual data" refers to information that includes image signals, and is the acquisition of visual elements such as still images and videos in digital format.
[0150] An "information terminal" is a device used to acquire auditory and visual data, and to process and transmit it.
[0151] A "data transmission path" refers to the means of communication or communication network used to transmit acquired data to an information processing device.
[0152] A "centralized information processing system" refers to a computer system for processing received data in an integrated manner, such as a server that provides central control.
[0153] An "automatic learning algorithm" is a method for automatically recognizing and analyzing data characteristics by learning certain rules or models based on received data.
[0154] "Support instructions" refer to information generated based on collected and analyzed data, including guidance and suggestions for actions to be taken by users and stakeholders.
[0155] "Priority" is an indicator used to determine the order in which to perform multiple actions or tasks, based on their importance and urgency.
[0156] This invention combines an information terminal that acquires audio and images with a central server system that processes this data, in order to provide effective support at disaster sites. The user uses the information terminal to acquire auditory and visual data consisting of audio and image signals. The information terminal incorporates a microphone and camera, enabling it to capture the user's voice and facial expressions.
[0157] The device transmits the acquired data to the server via Wi-Fi or mobile data communication. After transmission, the server analyzes the received data. Generative artificial intelligence models and machine learning algorithms are used for the analysis. Specifically, libraries such as TensorFlow are utilized to evaluate changes in voice tone and facial expressions. This process allows for the recognition of the user's emotional state and the situation at the site.
[0158] The server generates necessary support instructions from the analysis results, such as messages and encouraging information regarding psychological support. These instructions are created by a generative AI model using prompt sentences. An example of a prompt sentence is shown below.
[0159] "What kind of message should be generated when users are experiencing high stress levels?"
[0160] The generated instructions are sent back to the terminal, which displays them visually to the user and outputs them as voice guidance. For example, the terminal displays the message, "Observe the current situation calmly and take appropriate action," and provides the same guidance via voice. This entire process allows users to receive quick and appropriate support even in disaster areas, enabling them to act with a sense of security.
[0161] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0162] Step 1:
[0163] The device acquires the user's voice and facial expressions in the field. The input is the user's real-time voice and video, which is captured via a microphone and camera. Specifically, the device uses an acoustic sensor to pick up sound and a video sensor to record facial expressions. The output consists of digitized audio and image data.
[0164] Step 2:
[0165] The terminal transmits the acquired digitized data to the server. The input consists of digitized audio and image data, which are sent to the server via a data transmission path using a network communication module. Specifically, the terminal divides the data into packets and sends them to the server via a secure protocol. The output is the data stream received by the server.
[0166] Step 3:
[0167] The server analyzes the received data and extracts important information. The input is a data stream sent from the terminal, and it analyzes voice tone and facial expression patterns using a generative AI model and machine learning algorithms. Specifically, the server runs an emotion analysis model using TensorFlow and analyzes this data. The output is structured data indicating the user's emotional state.
[0168] Step 4:
[0169] The server generates appropriate support instructions based on the analysis results. The input is structured data indicating the user's emotional state, and it uses generative AI to generate support messages based on prompts. Specifically, the server generates instructions such as "Observe the current situation calmly and take appropriate action" through a generative model. The output is optimized support instructions.
[0170] Step 5:
[0171] The server sends the generated support instructions to the terminal. The input is the support instructions generated by the server, which are sent to the terminal via the data transmission path. Specifically, the server repackets the data and sends it using a secure protocol. The output is the support instruction data received by the terminal.
[0172] Step 6:
[0173] The terminal displays and audibly presents received support instructions to the user. The input is data of support instructions sent from the server, which is then communicated to the user via the display and speaker. Specifically, the terminal displays messages on the screen and plays instructions using speech synthesis technology. The output is information conveyed to the user visually and aurally.
[0174] (Application Example 2)
[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0176] In today's work environment, it is crucial to perform tasks efficiently while maintaining the health and safety of workers. However, conventional systems have struggled to monitor and respond to stress and fatigue levels in real time. Therefore, new methods are needed to ensure worker productivity and safety.
[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0178] In this invention, the server includes means for collecting environmental information of the work area, means for analyzing worker emotional data in real time and generating feedback, and means for generating and notifying instructions suggesting work breaks according to the emotional state. This makes it possible to create an environment that provides optimal work instructions while taking into account the worker's health condition.
[0179] "Voice data" refers to electromagnetic wave data obtained digitally from the worker's voice and used for analysis.
[0180] "Video data" refers to visual information obtained as digital images of a worker's facial expressions and movements, which is then used for analysis.
[0181] A "terminal device" is a device used to acquire audio and video data and transmit it to the main control unit.
[0182] An "information transmission network" is a communication infrastructure used to transmit data acquired by terminal devices to a main control unit.
[0183] The "main control unit" is a central processing unit that analyzes collected data, extracts knowledge, and generates instructions.
[0184] "Knowledge" refers to information about actions and responses that should be taken, extracted as a result of data analysis.
[0185] "Feedback" refers to advice and notifications provided to workers based on analyzed emotional data.
[0186] "Stress level" is a quantitative indicator of a worker's mental state, showing how it affects their normal work performance.
[0187] A "suggestion to pause work" is an instruction that recommends a temporary break to alleviate the stress levels of workers.
[0188] The system implementing this invention mainly consists of terminal equipment and a main control unit (server). The terminal equipment is responsible for acquiring audio and video data and transmitting this data to the server via an information transmission network. The terminal is capable of monitoring the surrounding environment and emotional state of the user, who is a worker, in real time.
[0189] The server analyzes received audio and video data to extract knowledge about the worker's stress level and emotional state. This analysis uses machine learning models (e.g., TensorFlow) to apply generated intelligent processing algorithms, enabling automatic data recognition and action decisions. Based on these results, the server generates emotionally appropriate work instructions and feedback, which are then sent to the terminal device.
[0190] For example, if the server determines from the worker's facial expression data that their stress level is high, it can immediately send a message to the terminal suggesting a short break. This is an important means of supporting the mental health of workers while maintaining their work efficiency. Furthermore, the generative AI model used is constantly updated and designed to flexibly adapt to new situations.
[0191] An example of a prompt might be: "Recognize the worker's fatigue level, analyze the data, and generate appropriate feedback. Specifically, be able to suggest breaks when stress levels are high." Based on this prompt, the system will take action to provide the worker with the best possible support.
[0192] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0193] Step 1:
[0194] The terminal acquires the worker's voice and video data in real time. Here, the terminal's camera and microphone are used to record the worker's facial expressions and voice as digital data. The input is raw data from the camera and microphone, which is temporarily stored on the terminal.
[0195] Step 2:
[0196] The terminal transmits collected audio and video data to the server using an information transmission network. The input is the audio and video data stored on the terminal, which is converted and transmitted using a transfer protocol. The output is the data stream received on the server.
[0197] Step 3:
[0198] The server begins analyzing the received data. Here, a generative AI model is used to analyze the audio and video data and assess the worker's emotional state and stress level. The input is a stream of audio and video data, and the output is the assessed emotional state. The analysis involves machine learning algorithms, and the data is processed within the model.
[0199] Step 4:
[0200] The server generates feedback or instructions based on the analysis results. In this process, it recommends the most appropriate action to the worker based on pre-configured prompts. The input is an assessment of emotional state, and the output is the generated feedback or instruction message. Specifically, it might generate suggestions such as, "Why don't you take a short break?"
[0201] Step 5:
[0202] The server sends feedback or instructions to the terminal. Here, the generated message is displayed on the terminal's screen and also provided as audio guidance. The input is the generated instruction message, and the output is the visual and auditory information received by the user (worker).
[0203] Step 6:
[0204] Users adjust their actions based on the feedback they receive. This reduces worker stress and maintains an efficient work environment. Input is instructions from the terminal, and output is the user's actual actions. Specific actions include taking suggested breaks.
[0205] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0206] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0207] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0208] [Second Embodiment]
[0209] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0210] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0211] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0212] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0213] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0215] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0216] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0217] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0218] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0219] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0220] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0221] This invention is a system that collects on-site environmental data using terminal devices that acquire acoustic and image signals, and transmits this data to a central processing unit via a communication network. The system includes the following program:
[0222] The terminal acquires acoustic and image signals and transmits this data to a server in real time. For example, a terminal carried by a rescue worker at a disaster site captures the surroundings and sounds through a camera and microphone mounted on their helmet. This allows the situation at the site to be acquired as digital data in real time.
[0223] The server analyzes the received acoustic and image signals. Using generative AI and machine learning algorithms, the server processes large amounts of data quickly and accurately, identifying damage and locations with a high probability of human casualties. This analysis allows for the extraction of information necessary to determine rescue priorities.
[0224] The server generates rescue orders from the extracted information, including guidelines on what kind of rescue operations should be carried out and where. For example, if it discovers a collapsed building from video data of a disaster area and, after analyzing the evacuation situation in the surrounding area, determines that there is a high probability that residents have not yet evacuated, it will issue an order for rapid rescue operations to that location. In this way, the server always generates accurate rescue orders based on the latest information.
[0225] Users receive these rescue instructions through their devices and carry out rescue operations based on them. The devices display the instructions sent from the server on the screen and notify the user via voice. This allows rescue workers to act efficiently and effectively. For example, users can check evacuation points and hazardous areas displayed on a map and act according to the instructions.
[0226] This system will enable faster and more efficient relief efforts during disasters, which is expected to reduce damage and improve safety.
[0227] The following describes the processing flow.
[0228] Step 1:
[0229] The device acquires audio and image signals.
[0230] The device is installed on-site, capturing surrounding video with its camera and recording ambient sounds with its microphone. This data is temporarily stored within the device.
[0231] Step 2:
[0232] The device sends data to the server.
[0233] The terminal transmits the acquired acoustic and image signals to the server in real time via the communication line. The data is encrypted to ensure the security of the transmission.
[0234] Step 3:
[0235] The server receives the data and begins analysis.
[0236] The server stores the received data in an analysis database and uses a generating AI to analyze the data. In this process, images are used to identify collapsed buildings and the presence or absence of survivors, and voices and rescue requests are extracted from audio.
[0237] Step 4:
[0238] The server extracts important information and generates rescue orders.
[0239] The server uses the analyzed data to assess the damage and evacuation situation in specific areas. Based on this information, it determines what kind of relief is needed in each area and creates specific relief orders.
[0240] Step 5:
[0241] The server sends a rescue order to the terminal.
[0242] The server sends the generated rescue instructions to the relevant terminals. This allows users receiving the instructions on-site to take immediate action.
[0243] Step 6:
[0244] The user checks and executes the rescue order on their device.
[0245] The user checks the instructions notified on their device and understands the details of the instructions from the on-screen interface. Then, they quickly carry out rescue operations based on the instructions.
[0246] (Example 1)
[0247] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0248] Disaster response activities on the ground require swift and accurate decision-making, but there is a lack of technology to collect reliable information in real time and provide timely instructions. Furthermore, relying on manual analysis of the collected information can lead to delays in response. This challenge poses a significant obstacle to minimizing damage and enhancing the effectiveness of relief efforts.
[0249] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0250] In this invention, the server includes means for collecting on-site environmental information using an information terminal that acquires sound and images, means for transmitting the acquired sound and images to a central processing unit via a communication path, and means for analyzing the received information in the central processing unit and extracting important content. This makes it possible to quickly calculate the priority of rescue efforts and generate and transmit optimal rescue guidelines.
[0251] "Acoustics" refers to signals based on sound waves generated from the environment and objects, and serves as a medium for obtaining information through these signals.
[0252] An "image" is digital data that represents visual information and is collected through cameras or other visual sensors.
[0253] An "information terminal" is a computer or electronic device used to acquire data such as sound and images and transmit it to other devices via a communication network.
[0254] A "communication path" is a path that includes the infrastructure and network technologies used to send and receive data.
[0255] A "central processing unit" is a computer system or server used to process and analyze received data.
[0256] "Generative artificial intelligence" is a technology or algorithm that uses machine learning based on large amounts of data to enable pattern recognition and automated analysis.
[0257] "Machine learning methods" are models and algorithms that automatically learn from data and perform inference and decision-making.
[0258] "Relief guidelines" are a set of action plans and instructions given to ensure that relief operations are carried out efficiently and effectively.
[0259] The system according to this invention consists mainly of three components: a terminal, a server, and a user.
[0260] First, the terminal is responsible for acquiring audio and video signals on-site. Specifically, a device equipped with a camera and microphone is used. This terminal continuously captures sound and video, and digitizes them in real time. This data is transmitted to a server via a communication network.
[0261] Next, the server analyzes the received data. The server possesses advanced computing power and uses generative AI models and machine learning algorithms to analyze the received acoustic and image signals in detail. This analysis process helps to understand the situation on site and identify high-risk areas and points requiring rescue. For example, the server can identify whether buildings have collapsed or to distinguish between different patterns of cries for help.
[0262] Finally, a key feature of this system is that it provides users with rescue guidelines based on the analysis results. The terminal receives instructions from the server, displays them to the user, and provides voice notifications as needed. The user can then quickly carry out rescue activities based on the information received. For example, they can move according to evacuation guidelines displayed by a map application.
[0263] This system provides rapid and accurate information processing in disaster response. A specific example of its use is inputting a prompt message to the server such as, "Analyze the latest acoustic and image data transmitted from the disaster area and identify locations where human casualties are likely." Based on this prompt message, the server performs the necessary analysis and generates results indicating the optimal relief strategy.
[0264] This invention provides a system that enables more efficient and rapid disaster response activities, thereby minimizing damage.
[0265] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0266] Step 1:
[0267] The terminal acquires acoustic and image signals from the site. It captures ambient sound and video as input and converts them into digital data. As output, it generates real-time aggregated digital acoustic and image data.
[0268] Step 2:
[0269] The terminal transmits the generated audio and image data over the communication network. In this step, it performs operations to properly packetize the digital data and send it to the server. The input is the data generated in step 1, and the output is the data transmission to the server.
[0270] Step 3:
[0271] The server analyzes the acoustic and image data received from the terminal. Using generative AI models and machine learning algorithms, it extracts important features from the input data. Specifically, it processes data to identify building collapses, human voices, and unusual sound patterns. As output, it generates information identifying high-risk areas and points requiring rescue.
[0272] Step 4:
[0273] The server generates rescue guidelines based on the information extracted in step 3 and sends them to the terminal. Based on the results of the generated AI, it forms action guidelines for high-priority rescue activities. Using the identified risk information as input, it creates specific rescue guidelines as output and sends them to the terminal.
[0274] Step 5:
[0275] Users initiate action using rescue guidelines received through their devices. The devices display the guidelines and provide audio alerts as needed. Input is rescue guidelines sent from the server, and output is information provided to the user through screen displays and audio notifications. Based on the information displayed on the devices, users quickly carry out evacuation and support activities.
[0276] (Application Example 1)
[0277] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0278] In modern society, in large-scale commercial facilities and event venues, a large number of people gather, and dangerous situations may occur. In conventional monitoring systems, there is a problem that real-time anomaly detection and prompt response are difficult, so it is necessary to solve these problems in a more efficient and effective way.
[0279] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0280] In this invention, the server includes means for collecting on-site environmental information using an information acquisition device that acquires acoustic signals and image signals, means for transmitting the acquired acoustic signals and image signals to a data processing device via a communication network, and in the data processing device, means for analyzing the received signals and extracting important information including abnormal events. Thereby, real-time anomaly detection and prompt response in commercial facilities and event venues become possible.
[0281] The "information acquisition device" is a device that acquires acoustic signals and image signals on-site and is used to collect environmental information.
[0282] The "communication network" is a communication system used to transmit signals acquired by the information acquisition device to the data processing device.
[0283] The "data processing device" is a device that analyzes the received acoustic signals and image signals and extracts important information.
[0284] The "abnormal event" is information indicating a situation deviating from the normal state or operation and is an event that should be particularly noted in the monitoring system.
[0285] The "important information" is information that is essential for real-time judgment and response among the data including abnormal events.
[0286] An "action instruction" is an instruction generated based on the extracted important information, which is displayed to the user to prompt specific actions.
[0287] The "user" refers to a person or organization that receives the action instruction sent to the information acquisition device and actually takes actions in the system of the present invention.
[0288] The system of the present invention provides a safe environment in commercial facilities and event venues by combining an information acquisition device, a communication network, and a data processing device. The information acquisition device uses cameras and microphones installed in smart glasses or smartphones to acquire acoustic signals and image signals in real time. These signals are transmitted to the data processing device via the communication network.
[0289] In the data processing device, generative artificial intelligence models such as TensorFlow and PyTorch built using Python are used to analyze the received acoustic signals and image signals. Through the analysis, abnormal events and important information in the environment are quickly extracted. For example, data processing is performed to detect abnormal behavior patterns and changes in volume.
[0290] The server generates appropriate action instructions based on these abnormal events and sends them to the information acquisition device. The information acquisition device displays the action instructions to the user and also uses voice output for notification. As a result, the user can grasp the situation in real time and respond quickly.
[0291] As a specific example, it is conceivable to utilize this system in large-scale events. A plurality of information acquisition devices are installed at the entrances and crowded areas of the event venue to monitor the movements of the visitors in real time. At this time, by setting an example of a prompt sentence such as "When a loud voice of a visitor is heard, it is judged as abnormal and notified to the security personnel" using the generative AI model, it becomes possible to immediately instruct countermeasures in case of an abnormality.
[0292] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0293] Step 1:
[0294] The terminal captures acoustic and image signals in real time using the camera and microphone of smart glasses or a smartphone to acquire on-site environmental information. The input is raw data from the site, and the output is digital signal data.
[0295] Step 2:
[0296] The terminal transmits the acquired acoustic and image signals as digital data to the data processing device via a communication network. The input is digitized acoustic and image data, and the output is communication data to the data processing device.
[0297] Step 3:
[0298] The server, in its data processing unit, runs a generative artificial intelligence model using TensorFlow or PyTorch to analyze received signals. The input is the transmitted signal data, and the output is anomalous events and important information extracted through the analysis. This analysis can, for example, detect peaks in audio waveforms or evaluate the degree of crowding from images.
[0299] Step 4:
[0300] The server generates appropriate action instructions based on the abnormal events obtained through analysis. The input is information about the abnormal events, and the output is specific action instructions. Using a generation AI model, instructions are generated according to prompt statements such as "If congestion is high in some areas, instruct security guards to check those areas."
[0301] Step 5:
[0302] The server transmits the generated action instructions to the information acquisition device, and the terminal notifies the user of these instructions. The input is action instruction data from the server, and the output is visual and auditory notifications to the user. The terminal displays the instructions on its screen and issues an audible warning.
[0303] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0304] The system according to the present invention not only collects on-site environmental data using a terminal device that acquires acoustic signals and image signals, but also combines an emotion engine that recognizes the user's emotion. This system is designed to realize effective rescue activities at the disaster site.
[0305] The terminal captures the voice and expression of the user in the disaster area and transmits this data to the server. As a result, the terminal can monitor the user's emotional state in real time and obtain emotion-based information as needed. For example, when a member of the rescue team is in a high-stress state, that information is recorded as emotion data.
[0306] The server analyzes the received emotion data and the environmental data acquired by the sensor. By utilizing generative artificial intelligence and machine learning algorithms, the server automatically recognizes the situation from these data and determines the actions to be taken accordingly. Emotion data is used to reevaluate the priority of rescue, especially in cases with a high stress level or urgency.
[0307] Based on this analysis result, the server generates an optimal rescue instruction and provides it to the user. For example, when the emotion engine highly evaluates the user's anxiety, the server transmits information on psychological support or an emergency encouragement message to the terminal. The user can receive this immediately and use it for the next action.
[0308] The terminal displays rescue instructions and emotion-based support information sent from the server to the user, and also provides voice guidance. Users can review the instructions and confidently take appropriate action. This system not only streamlines rescue operations but also contributes to maintaining the user's mental well-being.
[0309] Thus, the system of the present invention, which combines an emotion engine, is an important solution for supporting disaster relief activities in a multifaceted way and maintaining a safe and secure environment.
[0310] The following describes the processing flow.
[0311] Step 1:
[0312] The device acquires audio and image signals.
[0313] The terminals are deployed on-site to capture the user's voice and facial expressions in real time. Using cameras and microphones, they collect surrounding video and audio data, while simultaneously acquiring emotion-related data such as voice tone and changes in facial expressions.
[0314] Step 2:
[0315] The device analyzes the data using an emotion engine and sends the results to the server.
[0316] The emotion engine within the device determines the user's emotional state from collected acoustic and image signals. For example, it assesses stress levels from voice tone and speaking style, and detects signs of sadness or anxiety through facial expression analysis. The data, including these analysis results, is then transmitted to a server via the communication network.
[0317] Step 3:
[0318] The server analyzes all the data and extracts important information.
[0319] The server integrates and analyzes emotional and environmental data. Generative AI assesses the specific situation in the disaster area and determines which areas should be prioritized. By considering emotional data, it identifies the location and circumstances of users who particularly need psychological support.
[0320] Step 4:
[0321] The server generates optimal rescue instructions and psychological support information.
[0322] Based on the analysis results, the server determines the necessary relief measures and mental health support information. For example, the server might create instructions for a high-stress user, including resources for psychological counseling and encouraging messages.
[0323] Step 5:
[0324] The server generates instructions and sends them to the terminal.
[0325] The server sends rescue instructions and psychological support information to the terminal, enabling users to respond quickly. The data is encrypted and received instantly.
[0326] Step 6:
[0327] Users receive information and take action through their devices.
[0328] Users review the rescue instructions and support messages displayed on the device. This allows users to quickly take appropriate rescue actions while managing their emotional state. The device also provides voice guidance to help understand the instructions.
[0329] (Example 2)
[0330] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0331] In disaster areas, it is essential to grasp various emergency information and the emotional state of users in real time and to appropriately direct rescue operations. However, existing systems have problems in that they cannot adequately process information accurately and respond quickly while taking into account the emotional information of users. Therefore, it is necessary to develop a system that centrally processes diverse data, including emotional information, and provides more appropriate support instructions based on that information.
[0332] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0333] In this invention, the server includes means for collecting on-site situation information using an information terminal that acquires auditory and visual data, means for transmitting the acquired auditory and visual data to a centralized information processing device via a data transmission path, and means for analyzing the received data and identifying and extracting important elements in the information processing device. This enables rapid and accurate understanding of various situations at disaster sites and allows for appropriate support instructions that take into account the emotional state of users.
[0334] "Auditory data" refers to information that includes acoustic signals, such as speech and ambient sounds, acquired in digital format.
[0335] "Visual data" refers to information that includes image signals, and is the acquisition of visual elements such as still images and videos in digital format.
[0336] An "information terminal" is a device used to acquire auditory and visual data, and to process and transmit it.
[0337] A "data transmission path" refers to the means of communication or communication network used to transmit acquired data to an information processing device.
[0338] A "centralized information processing system" refers to a computer system for processing received data in an integrated manner, such as a server that provides central control.
[0339] An "automatic learning algorithm" is a method for automatically recognizing and analyzing data characteristics by learning certain rules or models based on received data.
[0340] "Support instructions" refer to information generated based on collected and analyzed data, including guidance and suggestions for actions to be taken by users and stakeholders.
[0341] "Priority" is an indicator used to determine the order in which to perform multiple actions or tasks, based on their importance and urgency.
[0342] This invention combines an information terminal that acquires audio and images with a central server system that processes this data, in order to provide effective support at disaster sites. The user uses the information terminal to acquire auditory and visual data consisting of audio and image signals. The information terminal incorporates a microphone and camera, enabling it to capture the user's voice and facial expressions.
[0343] The device transmits the acquired data to the server via Wi-Fi or mobile data communication. After transmission, the server analyzes the received data. Generative artificial intelligence models and machine learning algorithms are used for the analysis. Specifically, libraries such as TensorFlow are utilized to evaluate changes in voice tone and facial expressions. This process allows for the recognition of the user's emotional state and the situation at the site.
[0344] The server generates necessary support instructions from the analysis results, such as messages and encouraging information regarding psychological support. These instructions are created by a generative AI model using prompt sentences. An example of a prompt sentence is shown below.
[0345] "What kind of message should be generated when users are experiencing high stress levels?"
[0346] The generated instructions are sent back to the terminal, which displays them visually to the user and outputs them as voice guidance. For example, the terminal displays the message, "Observe the current situation calmly and take appropriate action," and provides the same guidance via voice. This entire process allows users to receive quick and appropriate support even in disaster areas, enabling them to act with a sense of security.
[0347] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0348] Step 1:
[0349] The device acquires the user's voice and facial expressions in the field. The input is the user's real-time voice and video, which is captured via a microphone and camera. Specifically, the device uses an acoustic sensor to pick up sound and a video sensor to record facial expressions. The output consists of digitized audio and image data.
[0350] Step 2:
[0351] The terminal transmits the acquired digitized data to the server. The input consists of digitized audio and image data, which are sent to the server via a data transmission path using a network communication module. Specifically, the terminal divides the data into packets and sends them to the server via a secure protocol. The output is the data stream received by the server.
[0352] Step 3:
[0353] The server analyzes the received data and extracts important information. The input is a data stream sent from the terminal, and it analyzes voice tone and facial expression patterns using a generative AI model and machine learning algorithms. Specifically, the server runs an emotion analysis model using TensorFlow and analyzes this data. The output is structured data indicating the user's emotional state.
[0354] Step 4:
[0355] The server generates appropriate support instructions based on the analysis results. The input is structured data indicating the user's emotional state, and it uses generative AI to generate support messages based on prompts. Specifically, the server generates instructions such as "Observe the current situation calmly and take appropriate action" through a generative model. The output is optimized support instructions.
[0356] Step 5:
[0357] The server sends the generated support instructions to the terminal. The input is the support instructions generated by the server, which are sent to the terminal via the data transmission path. Specifically, the server repackets the data and sends it using a secure protocol. The output is the support instruction data received by the terminal.
[0358] Step 6:
[0359] The terminal displays and audibly presents received support instructions to the user. The input is data of support instructions sent from the server, which is then communicated to the user via the display and speaker. Specifically, the terminal displays messages on the screen and plays instructions using speech synthesis technology. The output is information conveyed to the user visually and aurally.
[0360] (Application Example 2)
[0361] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0362] In today's work environment, it is crucial to perform tasks efficiently while maintaining the health and safety of workers. However, conventional systems have struggled to monitor and respond to stress and fatigue levels in real time. Therefore, new methods are needed to ensure worker productivity and safety.
[0363] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0364] In this invention, the server includes means for collecting environmental information of the work area, means for analyzing worker emotional data in real time and generating feedback, and means for generating and notifying instructions suggesting work breaks according to the emotional state. This makes it possible to create an environment that provides optimal work instructions while taking into account the worker's health condition.
[0365] "Voice data" refers to electromagnetic wave data obtained digitally from the worker's voice and used for analysis.
[0366] "Video data" refers to visual information obtained as digital images of a worker's facial expressions and movements, which is then used for analysis.
[0367] A "terminal device" is a device used to acquire audio and video data and transmit it to the main control unit.
[0368] An "information transmission network" is a communication infrastructure used to transmit data acquired by terminal devices to a main control unit.
[0369] The "main control unit" is a central processing unit that analyzes collected data, extracts knowledge, and generates instructions.
[0370] "Knowledge" refers to information about actions and responses that should be taken, extracted as a result of data analysis.
[0371] "Feedback" refers to advice and notifications provided to workers based on analyzed emotional data.
[0372] "Stress level" is a quantitative indicator of a worker's mental state, showing how it affects their normal work performance.
[0373] A "suggestion to pause work" is an instruction that recommends a temporary break to alleviate the stress levels of workers.
[0374] The system implementing this invention mainly consists of terminal equipment and a main control unit (server). The terminal equipment is responsible for acquiring audio and video data and transmitting this data to the server via an information transmission network. The terminal is capable of monitoring the surrounding environment and emotional state of the user, who is a worker, in real time.
[0375] The server analyzes received audio and video data to extract knowledge about the worker's stress level and emotional state. This analysis uses machine learning models (e.g., TensorFlow) to apply generated intelligent processing algorithms, enabling automatic data recognition and action decisions. Based on these results, the server generates emotionally appropriate work instructions and feedback, which are then sent to the terminal device.
[0376] For example, if the server determines from the worker's facial expression data that their stress level is high, it can immediately send a message to the terminal suggesting a short break. This is an important means of supporting the mental health of workers while maintaining their work efficiency. Furthermore, the generative AI model used is constantly updated and designed to flexibly adapt to new situations.
[0377] An example of a prompt might be: "Recognize the worker's fatigue level, analyze the data, and generate appropriate feedback. Specifically, be able to suggest breaks when stress levels are high." Based on this prompt, the system will take action to provide the worker with the best possible support.
[0378] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0379] Step 1:
[0380] The terminal acquires the worker's voice and video data in real time. Here, the terminal's camera and microphone are used to record the worker's facial expressions and voice as digital data. The input is raw data from the camera and microphone, which is temporarily stored on the terminal.
[0381] Step 2:
[0382] The terminal transmits collected audio and video data to the server using an information transmission network. The input is the audio and video data stored on the terminal, which is converted and transmitted using a transfer protocol. The output is the data stream received on the server.
[0383] Step 3:
[0384] The server begins analyzing the received data. Here, a generative AI model is used to analyze the audio and video data and assess the worker's emotional state and stress level. The input is a stream of audio and video data, and the output is the assessed emotional state. The analysis involves machine learning algorithms, and the data is processed within the model.
[0385] Step 4:
[0386] The server generates feedback or instructions based on the analysis results. In this process, it recommends the most appropriate action to the worker based on pre-configured prompts. The input is an assessment of emotional state, and the output is the generated feedback or instruction message. Specifically, it might generate suggestions such as, "Why don't you take a short break?"
[0387] Step 5:
[0388] The server sends feedback or instructions to the terminal. Here, the generated message is displayed on the terminal's screen and also provided as audio guidance. The input is the generated instruction message, and the output is the visual and auditory information received by the user (worker).
[0389] Step 6:
[0390] Users adjust their actions based on the feedback they receive. This reduces worker stress and maintains an efficient work environment. Input is instructions from the terminal, and output is the user's actual actions. Specific actions include taking suggested breaks.
[0391] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0392] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0393] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0394] [Third Embodiment]
[0395] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0396] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0397] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0398] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0399] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0400] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0401] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0402] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0403] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0404] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0405] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0406] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0407] This invention is a system that collects on-site environmental data using terminal devices that acquire acoustic and image signals, and transmits this data to a central processing unit via a communication network. The system includes the following program:
[0408] The terminal acquires acoustic and image signals and transmits this data to a server in real time. For example, a terminal carried by a rescue worker at a disaster site captures the surroundings and sounds through a camera and microphone mounted on their helmet. This allows the situation at the site to be acquired as digital data in real time.
[0409] The server analyzes the received acoustic and image signals. Using generative AI and machine learning algorithms, the server processes large amounts of data quickly and accurately, identifying damage and locations with a high probability of human casualties. This analysis allows for the extraction of information necessary to determine rescue priorities.
[0410] The server generates rescue orders from the extracted information, including guidelines on what kind of rescue operations should be carried out and where. For example, if it discovers a collapsed building from video data of a disaster area and, after analyzing the evacuation situation in the surrounding area, determines that there is a high probability that residents have not yet evacuated, it will issue an order for rapid rescue operations to that location. In this way, the server always generates accurate rescue orders based on the latest information.
[0411] Users receive these rescue instructions through their devices and carry out rescue operations based on them. The devices display the instructions sent from the server on the screen and notify the user via voice. This allows rescue workers to act efficiently and effectively. For example, users can check evacuation points and hazardous areas displayed on a map and act according to the instructions.
[0412] This system will enable faster and more efficient relief efforts during disasters, which is expected to reduce damage and improve safety.
[0413] The following describes the processing flow.
[0414] Step 1:
[0415] The device acquires audio and image signals.
[0416] The device is installed on-site, capturing surrounding video with its camera and recording ambient sounds with its microphone. This data is temporarily stored within the device.
[0417] Step 2:
[0418] The device sends data to the server.
[0419] The terminal transmits the acquired acoustic and image signals to the server in real time via the communication line. The data is encrypted to ensure the security of the transmission.
[0420] Step 3:
[0421] The server receives the data and begins analysis.
[0422] The server stores the received data in an analysis database and uses a generating AI to analyze the data. In this process, images are used to identify collapsed buildings and the presence or absence of survivors, and voices and rescue requests are extracted from audio.
[0423] Step 4:
[0424] The server extracts important information and generates rescue orders.
[0425] The server uses the analyzed data to assess the damage and evacuation situation in specific areas. Based on this information, it determines what kind of relief is needed in each area and creates specific relief orders.
[0426] Step 5:
[0427] The server sends a rescue order to the terminal.
[0428] The server sends the generated rescue instructions to the relevant terminals. This allows users receiving the instructions on-site to take immediate action.
[0429] Step 6:
[0430] The user checks and executes the rescue order on their device.
[0431] The user checks the instructions notified on their device and understands the details of the instructions from the on-screen interface. Then, they quickly carry out rescue operations based on the instructions.
[0432] (Example 1)
[0433] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0434] Disaster response activities on the ground require swift and accurate decision-making, but there is a lack of technology to collect reliable information in real time and provide timely instructions. Furthermore, relying on manual analysis of the collected information can lead to delays in response. This challenge poses a significant obstacle to minimizing damage and enhancing the effectiveness of relief efforts.
[0435] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0436] In this invention, the server includes means for collecting on-site environmental information using an information terminal that acquires sound and images, means for transmitting the acquired sound and images to a central processing unit via a communication path, and means for analyzing the received information in the central processing unit and extracting important content. This makes it possible to quickly calculate the priority of rescue efforts and generate and transmit optimal rescue guidelines.
[0437] "Acoustics" refers to signals based on sound waves generated from the environment and objects, and serves as a medium for obtaining information through these signals.
[0438] An "image" is digital data that represents visual information and is collected through cameras or other visual sensors.
[0439] An "information terminal" is a computer or electronic device used to acquire data such as sound and images and transmit it to other devices via a communication network.
[0440] A "communication path" is a path that includes the infrastructure and network technologies used to send and receive data.
[0441] A "central processing unit" is a computer system or server used to process and analyze received data.
[0442] "Generative artificial intelligence" is a technology or algorithm that uses machine learning based on large amounts of data to enable pattern recognition and automated analysis.
[0443] "Machine learning methods" are models and algorithms that automatically learn from data and perform inference and decision-making.
[0444] "Relief guidelines" are a set of action plans and instructions given to ensure that relief operations are carried out efficiently and effectively.
[0445] The system according to this invention consists mainly of three components: a terminal, a server, and a user.
[0446] First, the terminal is responsible for acquiring audio and video signals on-site. Specifically, a device equipped with a camera and microphone is used. This terminal continuously captures sound and video, and digitizes them in real time. This data is transmitted to a server via a communication network.
[0447] Next, the server analyzes the received data. The server possesses advanced computing power and uses generative AI models and machine learning algorithms to analyze the received acoustic and image signals in detail. This analysis process helps to understand the situation on site and identify high-risk areas and points requiring rescue. For example, the server can identify whether buildings have collapsed or to distinguish between different patterns of cries for help.
[0448] Finally, a key feature of this system is that it provides users with rescue guidelines based on the analysis results. The terminal receives instructions from the server, displays them to the user, and provides voice notifications as needed. The user can then quickly carry out rescue activities based on the information received. For example, they can move according to evacuation guidelines displayed by a map application.
[0449] This system provides rapid and accurate information processing in disaster response. A specific example of its use is inputting a prompt message to the server such as, "Analyze the latest acoustic and image data transmitted from the disaster area and identify locations where human casualties are likely." Based on this prompt message, the server performs the necessary analysis and generates results indicating the optimal relief strategy.
[0450] This invention provides a system that enables more efficient and rapid disaster response activities, thereby minimizing damage.
[0451] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0452] Step 1:
[0453] The terminal acquires acoustic and image signals from the site. It captures ambient sound and video as input and converts them into digital data. As output, it generates real-time aggregated digital acoustic and image data.
[0454] Step 2:
[0455] The terminal transmits the generated audio and image data over the communication network. In this step, it performs operations to properly packetize the digital data and send it to the server. The input is the data generated in step 1, and the output is the data transmission to the server.
[0456] Step 3:
[0457] The server analyzes the acoustic and image data received from the terminal. Using generative AI models and machine learning algorithms, it extracts important features from the input data. Specifically, it processes data to identify building collapses, human voices, and unusual sound patterns. As output, it generates information identifying high-risk areas and points requiring rescue.
[0458] Step 4:
[0459] The server generates rescue guidelines based on the information extracted in step 3 and sends them to the terminal. Based on the results of the generated AI, it forms action guidelines for high-priority rescue activities. Using the identified risk information as input, it creates specific rescue guidelines as output and sends them to the terminal.
[0460] Step 5:
[0461] Users initiate action using rescue guidelines received through their devices. The devices display the guidelines and provide audio alerts as needed. Input is rescue guidelines sent from the server, and output is information provided to the user through screen displays and audio notifications. Based on the information displayed on the devices, users quickly carry out evacuation and support activities.
[0462] (Application Example 1)
[0463] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0464] In modern society, large commercial facilities and event venues attract large crowds, potentially creating dangerous situations. Conventional monitoring systems struggle with real-time anomaly detection and rapid response, necessitating more efficient and effective solutions to these challenges.
[0465] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0466] In this invention, the server includes means for collecting on-site environmental information using an information acquisition device that acquires acoustic and image signals, means for transmitting the acquired acoustic and image signals to a data processing device via a communication network, and means for analyzing the received signals and extracting important information, including abnormal events, in the data processing device. This enables real-time anomaly detection and rapid response in commercial facilities and event venues.
[0467] An "information acquisition device" is a device that acquires acoustic and image signals on-site and is used to collect environmental information.
[0468] A "communication network" is a communication system used to transmit signals acquired from an information acquisition device to a data processing device.
[0469] A "data processing device" is a device that analyzes received acoustic and image signals and extracts important information.
[0470] An "abnormal event" is information indicating a situation that deviates from the normal state or operation, and is an event that requires particular attention in a monitoring system.
[0471] "Important information" refers to data, including abnormal events, that is essential for real-time decision-making and response.
[0472] "Action instructions" are instructions generated based on extracted key information, and are displayed to the user to prompt specific actions.
[0473] In the system of the present invention, "user" refers to a person or organization that receives action instructions transmitted to the information acquisition device and actually takes action.
[0474] The system of the present invention provides a secure environment in commercial facilities and event venues by combining an information acquisition device, a communication network, and a data processing device. The information acquisition device acquires acoustic and image signals in real time using a camera and microphone mounted on smart glasses or a smartphone. These signals are transmitted to the data processing device via the communication network.
[0475] The data processing unit uses generative artificial intelligence models such as TensorFlow and PyTorch, built with Python, to analyze received acoustic and image signals. This analysis allows for the rapid extraction of abnormal events and important information from the environment. For example, data processing is performed to detect abnormal behavioral patterns or changes in volume.
[0476] The server generates appropriate action instructions based on these abnormal events and transmits them to the information acquisition device. The information acquisition device displays the action instructions to the user and also provides notification via audio output. This enables the user to understand the situation in real time and respond quickly.
[0477] As a concrete example, this system could be used at large-scale events. Multiple information acquisition devices could be installed at the entrance of the event venue and in crowded areas to monitor the movement of attendees in real time. In this case, by using a generative AI model to set example prompts such as "If loud voices are heard from attendees, it will be judged as an anomaly and security guards will be notified," it would be possible to immediately instruct countermeasures in the event of an anomaly.
[0478] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0479] Step 1:
[0480] The terminal captures acoustic and image signals in real time using the camera and microphone of smart glasses or a smartphone to acquire on-site environmental information. The input is raw data from the site, and the output is digital signal data.
[0481] Step 2:
[0482] The terminal transmits the acquired acoustic and image signals as digital data to the data processing device via a communication network. The input is digitized acoustic and image data, and the output is communication data to the data processing device.
[0483] Step 3:
[0484] The server, in its data processing unit, runs a generative artificial intelligence model using TensorFlow or PyTorch to analyze received signals. The input is the transmitted signal data, and the output is anomalous events and important information extracted through the analysis. This analysis can, for example, detect peaks in audio waveforms or evaluate the degree of crowding from images.
[0485] Step 4:
[0486] The server generates appropriate action instructions based on the abnormal events obtained through analysis. The input is information about the abnormal events, and the output is specific action instructions. Using a generation AI model, instructions are generated according to prompt statements such as "If congestion is high in some areas, instruct security guards to check those areas."
[0487] Step 5:
[0488] The server transmits the generated action instructions to the information acquisition device, and the terminal notifies the user of these instructions. The input is action instruction data from the server, and the output is visual and auditory notification to the user. The terminal displays the instructions on its screen and issues an audible warning.
[0489] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0490] The system in this invention not only collects on-site environmental data using terminal devices that acquire acoustic and image signals, but also incorporates an emotion engine that recognizes the user's emotions. This system is designed to enable effective rescue operations at disaster sites.
[0491] The device captures the voice and facial expressions of users in the disaster area and sends this data to a server. This allows the device to monitor the user's emotional state in real time and obtain emotion-based information as needed. For example, if a member of a rescue team is experiencing high stress levels, this information is recorded as emotional data.
[0492] The server analyzes received emotional data and environmental data acquired by sensors. By utilizing generative artificial intelligence and machine learning algorithms, the server automatically recognizes the situation from this data and determines the appropriate course of action. Emotional data is used to re-evaluate the priority of rescue efforts, particularly in cases of high stress levels or urgency.
[0493] Based on these analysis results, the server generates and provides the user with the most appropriate rescue instructions. For example, if the emotion engine highly values the user's anxiety, the server sends information about psychological support and urgent encouraging messages to the device. The user receives this immediately and can use it to guide their next actions.
[0494] The terminal displays rescue instructions and emotion-based support information sent from the server to the user, and also provides voice guidance. Users can review the instructions and confidently take appropriate action. This system not only streamlines rescue operations but also contributes to maintaining the user's mental well-being.
[0495] Thus, the system of the present invention, which combines an emotion engine, is an important solution for supporting disaster relief activities in a multifaceted way and maintaining a safe and secure environment.
[0496] The following describes the processing flow.
[0497] Step 1:
[0498] The device acquires audio and image signals.
[0499] The terminals are deployed on-site to capture the user's voice and facial expressions in real time. Using cameras and microphones, they collect surrounding video and audio data, while simultaneously acquiring emotion-related data such as voice tone and changes in facial expressions.
[0500] Step 2:
[0501] The device analyzes the data using an emotion engine and sends the results to the server.
[0502] The emotion engine within the device determines the user's emotional state from collected acoustic and image signals. For example, it assesses stress levels from voice tone and speaking style, and detects signs of sadness or anxiety through facial expression analysis. The data, including these analysis results, is then transmitted to a server via the communication network.
[0503] Step 3:
[0504] The server analyzes all the data and extracts important information.
[0505] The server integrates and analyzes emotional and environmental data. Generative AI assesses the specific situation in the disaster area and determines which areas should be prioritized. By considering emotional data, it identifies the location and circumstances of users who particularly need psychological support.
[0506] Step 4:
[0507] The server generates optimal rescue instructions and psychological support information.
[0508] Based on the analysis results, the server determines the necessary relief measures and mental health support information. For example, the server might create instructions for a high-stress user, including resources for psychological counseling and encouraging messages.
[0509] Step 5:
[0510] The server generates instructions and sends them to the terminal.
[0511] The server sends rescue instructions and psychological support information to the terminal, enabling users to respond quickly. The data is encrypted and received instantly.
[0512] Step 6:
[0513] Users receive information and take action through their devices.
[0514] Users review the rescue instructions and support messages displayed on the device. This allows users to quickly take appropriate rescue actions while managing their emotional state. The device also provides voice guidance to help understand the instructions.
[0515] (Example 2)
[0516] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0517] In disaster areas, it is essential to grasp various emergency information and the emotional state of users in real time and to appropriately direct rescue operations. However, existing systems have problems in that they cannot adequately process information accurately and respond quickly while taking into account the emotional information of users. Therefore, it is necessary to develop a system that centrally processes diverse data, including emotional information, and provides more appropriate support instructions based on that information.
[0518] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0519] In this invention, the server includes means for collecting on-site situation information using an information terminal that acquires auditory and visual data, means for transmitting the acquired auditory and visual data to a centralized information processing device via a data transmission path, and means for analyzing the received data and identifying and extracting important elements in the information processing device. This enables rapid and accurate understanding of various situations at disaster sites and allows for appropriate support instructions that take into account the emotional state of users.
[0520] "Auditory data" refers to information that includes acoustic signals, such as speech and ambient sounds, acquired in digital format.
[0521] "Visual data" refers to information that includes image signals, and is the acquisition of visual elements such as still images and videos in digital format.
[0522] An "information terminal" is a device used to acquire auditory and visual data, and to process and transmit it.
[0523] A "data transmission path" refers to the means of communication or communication network used to transmit acquired data to an information processing device.
[0524] A "centralized information processing system" refers to a computer system for processing received data in an integrated manner, such as a server that provides central control.
[0525] An "automatic learning algorithm" is a method for automatically recognizing and analyzing data characteristics by learning certain rules or models based on received data.
[0526] "Support instructions" refer to information generated based on collected and analyzed data, including guidance and suggestions for actions to be taken by users and stakeholders.
[0527] "Priority" is an indicator used to determine the order in which to perform multiple actions or tasks, based on their importance and urgency.
[0528] This invention combines an information terminal that acquires audio and images with a central server system that processes this data, in order to provide effective support at disaster sites. The user uses the information terminal to acquire auditory and visual data consisting of audio and image signals. The information terminal incorporates a microphone and camera, enabling it to capture the user's voice and facial expressions.
[0529] The device transmits the acquired data to the server via Wi-Fi or mobile data communication. After transmission, the server analyzes the received data. Generative artificial intelligence models and machine learning algorithms are used for the analysis. Specifically, libraries such as TensorFlow are utilized to evaluate changes in voice tone and facial expressions. This process allows for the recognition of the user's emotional state and the situation at the site.
[0530] The server generates necessary support instructions from the analysis results, such as messages and encouraging information regarding psychological support. These instructions are created by a generative AI model using prompt sentences. An example of a prompt sentence is shown below.
[0531] "What kind of message should be generated when users are experiencing high stress levels?"
[0532] The generated instructions are sent back to the terminal, which displays them visually to the user and outputs them as voice guidance. For example, the terminal displays the message, "Observe the current situation calmly and take appropriate action," and provides the same guidance via voice. This entire process allows users to receive quick and appropriate support even in disaster areas, enabling them to act with a sense of security.
[0533] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0534] Step 1:
[0535] The device acquires the user's voice and facial expressions in the field. The input is the user's real-time voice and video, which is captured via a microphone and camera. Specifically, the device uses an acoustic sensor to pick up sound and a video sensor to record facial expressions. The output consists of digitized audio and image data.
[0536] Step 2:
[0537] The terminal transmits the acquired digitized data to the server. The input consists of digitized audio and image data, which are sent to the server via a data transmission path using a network communication module. Specifically, the terminal divides the data into packets and sends them to the server via a secure protocol. The output is the data stream received by the server.
[0538] Step 3:
[0539] The server analyzes the received data and extracts important information. The input is a data stream sent from the terminal, and it analyzes voice tone and facial expression patterns using a generative AI model and machine learning algorithms. Specifically, the server runs an emotion analysis model using TensorFlow and analyzes this data. The output is structured data indicating the user's emotional state.
[0540] Step 4:
[0541] The server generates appropriate support instructions based on the analysis results. The input is structured data indicating the user's emotional state, and it uses generative AI to generate support messages based on prompts. Specifically, the server generates instructions such as "Observe the current situation calmly and take appropriate action" through a generative model. The output is optimized support instructions.
[0542] Step 5:
[0543] The server sends the generated support instructions to the terminal. The input is the support instructions generated by the server, which are sent to the terminal via the data transmission path. Specifically, the server repackets the data and sends it using a secure protocol. The output is the support instruction data received by the terminal.
[0544] Step 6:
[0545] The terminal displays and audibly presents received support instructions to the user. The input is data of support instructions sent from the server, which is then communicated to the user via the display and speaker. Specifically, the terminal displays messages on the screen and plays instructions using speech synthesis technology. The output is information conveyed to the user visually and aurally.
[0546] (Application Example 2)
[0547] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0548] In today's work environment, it is crucial to perform tasks efficiently while maintaining the health and safety of workers. However, conventional systems have struggled to monitor and respond to stress and fatigue levels in real time. Therefore, new methods are needed to ensure worker productivity and safety.
[0549] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0550] In this invention, the server includes means for collecting environmental information of the work area, means for analyzing worker emotional data in real time and generating feedback, and means for generating and notifying instructions suggesting work breaks according to the emotional state. This makes it possible to create an environment that provides optimal work instructions while taking into account the worker's health condition.
[0551] "Voice data" refers to electromagnetic wave data obtained digitally from the worker's voice and used for analysis.
[0552] "Video data" refers to visual information obtained as digital images of a worker's facial expressions and movements, which is then used for analysis.
[0553] A "terminal device" is a device used to acquire audio and video data and transmit it to the main control unit.
[0554] An "information transmission network" is a communication infrastructure used to transmit data acquired by terminal devices to a main control unit.
[0555] The "main control unit" is a central processing unit that analyzes collected data, extracts knowledge, and generates instructions.
[0556] "Knowledge" refers to information about actions and responses that should be taken, extracted as a result of data analysis.
[0557] "Feedback" refers to advice and notifications provided to workers based on analyzed emotional data.
[0558] "Stress level" is a quantitative indicator of a worker's mental state, showing how it affects their normal work performance.
[0559] A "suggestion to pause work" is an instruction that recommends a temporary break to alleviate the stress levels of workers.
[0560] The system implementing this invention mainly consists of terminal equipment and a main control unit (server). The terminal equipment is responsible for acquiring audio and video data and transmitting this data to the server via an information transmission network. The terminal is capable of monitoring the surrounding environment and emotional state of the user, who is a worker, in real time.
[0561] The server analyzes received audio and video data to extract knowledge about the worker's stress level and emotional state. This analysis uses machine learning models (e.g., TensorFlow) to apply generated intelligent processing algorithms, enabling automatic data recognition and action decisions. Based on these results, the server generates emotionally appropriate work instructions and feedback, which are then sent to the terminal device.
[0562] For example, if the server determines from the worker's facial expression data that their stress level is high, it can immediately send a message to the terminal suggesting a short break. This is an important means of supporting the mental health of workers while maintaining their work efficiency. Furthermore, the generative AI model used is constantly updated and designed to flexibly adapt to new situations.
[0563] An example of a prompt might be: "Recognize the worker's fatigue level, analyze the data, and generate appropriate feedback. Specifically, be able to suggest breaks when stress levels are high." Based on this prompt, the system will take action to provide the worker with the best possible support.
[0564] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0565] Step 1:
[0566] The terminal acquires the worker's voice and video data in real time. Here, the terminal's camera and microphone are used to record the worker's facial expressions and voice as digital data. The input is raw data from the camera and microphone, which is temporarily stored on the terminal.
[0567] Step 2:
[0568] The terminal transmits collected audio and video data to the server using an information transmission network. The input is the audio and video data stored on the terminal, which is converted and transmitted using a transfer protocol. The output is the data stream received on the server.
[0569] Step 3:
[0570] The server begins analyzing the received data. Here, a generative AI model is used to analyze the audio and video data and assess the worker's emotional state and stress level. The input is a stream of audio and video data, and the output is the assessed emotional state. The analysis involves machine learning algorithms, and the data is processed within the model.
[0571] Step 4:
[0572] The server generates feedback or instructions based on the analysis results. In this process, it recommends the most appropriate action to the worker based on pre-configured prompts. The input is an assessment of emotional state, and the output is the generated feedback or instruction message. Specifically, it might generate suggestions such as, "Why don't you take a short break?"
[0573] Step 5:
[0574] The server sends feedback or instructions to the terminal. Here, the generated message is displayed on the terminal's screen and also provided as audio guidance. The input is the generated instruction message, and the output is the visual and auditory information received by the user (worker).
[0575] Step 6:
[0576] Users adjust their actions based on the feedback they receive. This reduces worker stress and maintains an efficient work environment. Input is instructions from the terminal, and output is the user's actual actions. Specific actions include taking suggested breaks.
[0577] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0578] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0579] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0580] [Fourth Embodiment]
[0581] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0582] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0583] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0584] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0585] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0586] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0587] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0588] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0589] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0590] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0591] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0592] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0593] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0594] This invention is a system that collects on-site environmental data using terminal devices that acquire acoustic and image signals, and transmits this data to a central processing unit via a communication network. The system includes the following program:
[0595] The terminal acquires acoustic and image signals and transmits this data to a server in real time. For example, a terminal carried by a rescue worker at a disaster site captures the surroundings and sounds through a camera and microphone mounted on their helmet. This allows the situation at the site to be acquired as digital data in real time.
[0596] The server analyzes the received acoustic and image signals. Using generative AI and machine learning algorithms, the server processes large amounts of data quickly and accurately, identifying damage and locations with a high probability of human casualties. This analysis allows for the extraction of information necessary to determine rescue priorities.
[0597] The server generates rescue orders from the extracted information, including guidelines on what kind of rescue operations should be carried out and where. For example, if it discovers a collapsed building from video data of a disaster area and, after analyzing the evacuation situation in the surrounding area, determines that there is a high probability that residents have not yet evacuated, it will issue an order for rapid rescue operations to that location. In this way, the server always generates accurate rescue orders based on the latest information.
[0598] Users receive these rescue instructions through their devices and carry out rescue operations based on them. The devices display the instructions sent from the server on the screen and notify the user via voice. This allows rescue workers to act efficiently and effectively. For example, users can check evacuation points and hazardous areas displayed on a map and act according to the instructions.
[0599] This system will enable faster and more efficient relief efforts during disasters, which is expected to reduce damage and improve safety.
[0600] The following describes the processing flow.
[0601] Step 1:
[0602] The device acquires audio and image signals.
[0603] The device is installed on-site, capturing surrounding video with its camera and recording ambient sounds with its microphone. This data is temporarily stored within the device.
[0604] Step 2:
[0605] The device sends data to the server.
[0606] The terminal transmits the acquired acoustic and image signals to the server in real time via the communication line. The data is encrypted to ensure the security of the transmission.
[0607] Step 3:
[0608] The server receives the data and begins analysis.
[0609] The server stores the received data in an analysis database and uses a generating AI to analyze the data. In this process, images are used to identify collapsed buildings and the presence or absence of survivors, and voices and rescue requests are extracted from audio.
[0610] Step 4:
[0611] The server extracts important information and generates rescue orders.
[0612] The server uses the analyzed data to assess the damage and evacuation situation in specific areas. Based on this information, it determines what kind of relief is needed in each area and creates specific relief orders.
[0613] Step 5:
[0614] The server sends a rescue order to the terminal.
[0615] The server sends the generated rescue instructions to the relevant terminals. This allows users receiving the instructions on-site to take immediate action.
[0616] Step 6:
[0617] The user checks and executes the rescue order on their device.
[0618] The user checks the instructions notified on their device and understands the details of the instructions from the on-screen interface. Then, they quickly carry out rescue operations based on the instructions.
[0619] (Example 1)
[0620] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0621] Disaster response activities on the ground require swift and accurate decision-making, but there is a lack of technology to collect reliable information in real time and provide timely instructions. Furthermore, relying on manual analysis of the collected information can lead to delays in response. This challenge poses a significant obstacle to minimizing damage and enhancing the effectiveness of relief efforts.
[0622] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0623] In this invention, the server includes means for collecting on-site environmental information using an information terminal that acquires sound and images, means for transmitting the acquired sound and images to a central processing unit via a communication path, and means for analyzing the received information in the central processing unit and extracting important content. This makes it possible to quickly calculate the priority of rescue efforts and generate and transmit optimal rescue guidelines.
[0624] "Acoustics" refers to signals based on sound waves generated from the environment and objects, and serves as a medium for obtaining information through these signals.
[0625] An "image" is digital data that represents visual information and is collected through cameras or other visual sensors.
[0626] An "information terminal" is a computer or electronic device used to acquire data such as sound and images and transmit it to other devices via a communication network.
[0627] A "communication path" is a path that includes the infrastructure and network technologies used to send and receive data.
[0628] A "central processing unit" is a computer system or server used to process and analyze received data.
[0629] "Generative artificial intelligence" is a technology or algorithm that uses machine learning based on large amounts of data to enable pattern recognition and automated analysis.
[0630] "Machine learning methods" are models and algorithms that automatically learn from data and perform inference and decision-making.
[0631] "Relief guidelines" are a set of action plans and instructions given to ensure that relief operations are carried out efficiently and effectively.
[0632] The system according to this invention consists mainly of three components: a terminal, a server, and a user.
[0633] First, the terminal is responsible for acquiring audio and video signals on-site. Specifically, a device equipped with a camera and microphone is used. This terminal continuously captures sound and video, and digitizes them in real time. This data is transmitted to a server via a communication network.
[0634] Next, the server analyzes the received data. The server possesses advanced computing power and uses generative AI models and machine learning algorithms to analyze the received acoustic and image signals in detail. This analysis process helps to understand the situation on site and identify high-risk areas and points requiring rescue. For example, the server can identify whether buildings have collapsed or to distinguish between different patterns of cries for help.
[0635] Finally, a key feature of this system is that it provides users with rescue guidelines based on the analysis results. The terminal receives instructions from the server, displays them to the user, and provides voice notifications as needed. The user can then quickly carry out rescue activities based on the information received. For example, they can move according to evacuation guidelines displayed by a map application.
[0636] This system provides rapid and accurate information processing in disaster response. A specific example of its use is inputting a prompt message to the server such as, "Analyze the latest acoustic and image data transmitted from the disaster area and identify locations where human casualties are likely." Based on this prompt message, the server performs the necessary analysis and generates results indicating the optimal relief strategy.
[0637] This invention provides a system that enables more efficient and rapid disaster response activities, thereby minimizing damage.
[0638] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0639] Step 1:
[0640] The terminal acquires acoustic and image signals from the site. It captures ambient sound and video as input and converts them into digital data. As output, it generates real-time aggregated digital acoustic and image data.
[0641] Step 2:
[0642] The terminal transmits the generated audio and image data over the communication network. In this step, it performs operations to properly packetize the digital data and send it to the server. The input is the data generated in step 1, and the output is the data transmission to the server.
[0643] Step 3:
[0644] The server analyzes the acoustic and image data received from the terminal. Using generative AI models and machine learning algorithms, it extracts important features from the input data. Specifically, it processes data to identify building collapses, human voices, and unusual sound patterns. As output, it generates information identifying high-risk areas and points requiring rescue.
[0645] Step 4:
[0646] The server generates rescue guidelines based on the information extracted in step 3 and sends them to the terminal. Based on the results of the generated AI, it forms action guidelines for high-priority rescue activities. Using the identified risk information as input, it creates specific rescue guidelines as output and sends them to the terminal.
[0647] Step 5:
[0648] Users initiate action using rescue guidelines received through their devices. The devices display the guidelines and provide audio alerts as needed. Input is rescue guidelines sent from the server, and output is information provided to the user through screen displays and audio notifications. Based on the information displayed on the devices, users quickly carry out evacuation and support activities.
[0649] (Application Example 1)
[0650] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0651] In modern society, large commercial facilities and event venues attract large crowds, potentially creating dangerous situations. Conventional monitoring systems struggle with real-time anomaly detection and rapid response, necessitating more efficient and effective solutions to these challenges.
[0652] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0653] In this invention, the server includes means for collecting on-site environmental information using an information acquisition device that acquires acoustic and image signals, means for transmitting the acquired acoustic and image signals to a data processing device via a communication network, and means for analyzing the received signals and extracting important information, including abnormal events, in the data processing device. This enables real-time anomaly detection and rapid response in commercial facilities and event venues.
[0654] An "information acquisition device" is a device that acquires acoustic and image signals on-site and is used to collect environmental information.
[0655] A "communication network" is a communication system used to transmit signals acquired from an information acquisition device to a data processing device.
[0656] A "data processing device" is a device that analyzes received acoustic and image signals and extracts important information.
[0657] An "abnormal event" is information indicating a situation that deviates from the normal state or operation, and is an event that requires particular attention in a monitoring system.
[0658] "Important information" refers to data, including abnormal events, that is essential for real-time decision-making and response.
[0659] "Action instructions" are instructions generated based on extracted key information, and are displayed to the user to prompt specific actions.
[0660] In the system of the present invention, "user" refers to a person or organization that receives action instructions transmitted to the information acquisition device and actually takes action.
[0661] The system of the present invention provides a secure environment in commercial facilities and event venues by combining an information acquisition device, a communication network, and a data processing device. The information acquisition device acquires acoustic and image signals in real time using a camera and microphone mounted on smart glasses or a smartphone. These signals are transmitted to the data processing device via the communication network.
[0662] The data processing unit uses generative artificial intelligence models such as TensorFlow and PyTorch, built with Python, to analyze received acoustic and image signals. This analysis allows for the rapid extraction of abnormal events and important information from the environment. For example, data processing is performed to detect abnormal behavioral patterns or changes in volume.
[0663] The server generates appropriate action instructions based on these abnormal events and transmits them to the information acquisition device. The information acquisition device displays the action instructions to the user and also provides notification via audio output. This enables the user to understand the situation in real time and respond quickly.
[0664] As a concrete example, this system could be used at large-scale events. Multiple information acquisition devices could be installed at the entrance of the event venue and in crowded areas to monitor the movement of attendees in real time. In this case, by using a generative AI model to set example prompts such as "If loud voices are heard from attendees, it will be judged as an anomaly and security guards will be notified," it would be possible to immediately instruct countermeasures in the event of an anomaly.
[0665] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0666] Step 1:
[0667] The terminal captures acoustic and image signals in real time using the camera and microphone of smart glasses or a smartphone to acquire on-site environmental information. The input is raw data from the site, and the output is digital signal data.
[0668] Step 2:
[0669] The terminal transmits the acquired acoustic and image signals as digital data to the data processing device via a communication network. The input is digitized acoustic and image data, and the output is communication data to the data processing device.
[0670] Step 3:
[0671] The server, in its data processing unit, runs a generative artificial intelligence model using TensorFlow or PyTorch to analyze received signals. The input is the transmitted signal data, and the output is anomalous events and important information extracted through the analysis. This analysis can, for example, detect peaks in audio waveforms or evaluate the degree of crowding from images.
[0672] Step 4:
[0673] The server generates appropriate action instructions based on the abnormal events obtained through analysis. The input is information about the abnormal events, and the output is specific action instructions. Using a generation AI model, instructions are generated according to prompt statements such as "If congestion is high in some areas, instruct security guards to check those areas."
[0674] Step 5:
[0675] The server transmits the generated action instructions to the information acquisition device, and the terminal notifies the user of these instructions. The input is action instruction data from the server, and the output is visual and auditory notification to the user. The terminal displays the instructions on its screen and issues an audible warning.
[0676] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0677] The system in this invention not only collects on-site environmental data using terminal devices that acquire acoustic and image signals, but also incorporates an emotion engine that recognizes the user's emotions. This system is designed to enable effective rescue operations at disaster sites.
[0678] The device captures the voice and facial expressions of users in the disaster area and sends this data to a server. This allows the device to monitor the user's emotional state in real time and obtain emotion-based information as needed. For example, if a member of a rescue team is experiencing high stress levels, this information is recorded as emotional data.
[0679] The server analyzes received emotional data and environmental data acquired by sensors. By utilizing generative artificial intelligence and machine learning algorithms, the server automatically recognizes the situation from this data and determines the appropriate course of action. Emotional data is used to re-evaluate the priority of rescue efforts, particularly in cases of high stress levels or urgency.
[0680] Based on these analysis results, the server generates and provides the user with the most appropriate rescue instructions. For example, if the emotion engine highly values the user's anxiety, the server sends information about psychological support and urgent encouraging messages to the device. The user receives this immediately and can use it to guide their next actions.
[0681] The terminal displays rescue instructions and emotion-based support information sent from the server to the user, and also provides voice guidance. Users can review the instructions and confidently take appropriate action. This system not only streamlines rescue operations but also contributes to maintaining the user's mental well-being.
[0682] Thus, the system of the present invention, which combines an emotion engine, is an important solution for supporting disaster relief activities in a multifaceted way and maintaining a safe and secure environment.
[0683] The following describes the processing flow.
[0684] Step 1:
[0685] The device acquires audio and image signals.
[0686] The terminals are deployed on-site to capture the user's voice and facial expressions in real time. Using cameras and microphones, they collect surrounding video and audio data, while simultaneously acquiring emotion-related data such as voice tone and changes in facial expressions.
[0687] Step 2:
[0688] The device analyzes the data using an emotion engine and sends the results to the server.
[0689] The emotion engine within the device determines the user's emotional state from collected acoustic and image signals. For example, it assesses stress levels from voice tone and speaking style, and detects signs of sadness or anxiety through facial expression analysis. The data, including these analysis results, is then transmitted to a server via the communication network.
[0690] Step 3:
[0691] The server analyzes all the data and extracts important information.
[0692] The server integrates and analyzes emotional and environmental data. Generative AI assesses the specific situation in the disaster area and determines which areas should be prioritized. By considering emotional data, it identifies the location and circumstances of users who particularly need psychological support.
[0693] Step 4:
[0694] The server generates optimal rescue instructions and psychological support information.
[0695] Based on the analysis results, the server determines the necessary relief measures and mental health support information. For example, the server might create instructions for a high-stress user, including resources for psychological counseling and encouraging messages.
[0696] Step 5:
[0697] The server generates instructions and sends them to the terminal.
[0698] The server sends rescue instructions and psychological support information to the terminal, enabling users to respond quickly. The data is encrypted and received instantly.
[0699] Step 6:
[0700] Users receive information and take action through their devices.
[0701] Users review the rescue instructions and support messages displayed on the device. This allows users to quickly take appropriate rescue actions while managing their emotional state. The device also provides voice guidance to help understand the instructions.
[0702] (Example 2)
[0703] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0704] In disaster areas, it is essential to grasp various emergency information and the emotional state of users in real time and to appropriately direct rescue operations. However, existing systems have problems in that they cannot adequately process information accurately and respond quickly while taking into account the emotional information of users. Therefore, it is necessary to develop a system that centrally processes diverse data, including emotional information, and provides more appropriate support instructions based on that information.
[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0706] In this invention, the server includes means for collecting on-site situation information using an information terminal that acquires auditory and visual data, means for transmitting the acquired auditory and visual data to a centralized information processing device via a data transmission path, and means for analyzing the received data and identifying and extracting important elements in the information processing device. This enables rapid and accurate understanding of various situations at disaster sites and allows for appropriate support instructions that take into account the emotional state of users.
[0707] "Auditory data" refers to information that includes acoustic signals, such as speech and ambient sounds, acquired in digital format.
[0708] "Visual data" refers to information that includes image signals, and is the acquisition of visual elements such as still images and videos in digital format.
[0709] An "information terminal" is a device used to acquire auditory and visual data, and to process and transmit it.
[0710] A "data transmission path" refers to the means of communication or communication network used to transmit acquired data to an information processing device.
[0711] A "centralized information processing system" refers to a computer system for processing received data in an integrated manner, such as a server that provides central control.
[0712] An "automatic learning algorithm" is a method for automatically recognizing and analyzing data characteristics by learning certain rules or models based on received data.
[0713] "Support instructions" refer to information generated based on collected and analyzed data, including guidance and suggestions for actions to be taken by users and stakeholders.
[0714] "Priority" is an indicator used to determine the order in which to perform multiple actions or tasks, based on their importance and urgency.
[0715] This invention combines an information terminal that acquires audio and images with a central server system that processes this data, in order to provide effective support at disaster sites. The user uses the information terminal to acquire auditory and visual data consisting of audio and image signals. The information terminal incorporates a microphone and camera, enabling it to capture the user's voice and facial expressions.
[0716] The device transmits the acquired data to the server via Wi-Fi or mobile data communication. After transmission, the server analyzes the received data. Generative artificial intelligence models and machine learning algorithms are used for the analysis. Specifically, libraries such as TensorFlow are utilized to evaluate changes in voice tone and facial expressions. This process allows for the recognition of the user's emotional state and the situation at the site.
[0717] The server generates necessary support instructions from the analysis results, such as messages and encouraging information regarding psychological support. These instructions are created by a generative AI model using prompt sentences. An example of a prompt sentence is shown below.
[0718] "What kind of message should be generated when users are experiencing high stress levels?"
[0719] The generated instructions are sent back to the terminal, which displays them visually to the user and outputs them as voice guidance. For example, the terminal displays the message, "Observe the current situation calmly and take appropriate action," and provides the same guidance via voice. This entire process allows users to receive quick and appropriate support even in disaster areas, enabling them to act with a sense of security.
[0720] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0721] Step 1:
[0722] The device acquires the user's voice and facial expressions in the field. The input is the user's real-time voice and video, which is captured via a microphone and camera. Specifically, the device uses an acoustic sensor to pick up sound and a video sensor to record facial expressions. The output consists of digitized audio and image data.
[0723] Step 2:
[0724] The terminal transmits the acquired digitized data to the server. The input consists of digitized audio and image data, which are sent to the server via a data transmission path using a network communication module. Specifically, the terminal divides the data into packets and sends them to the server via a secure protocol. The output is the data stream received by the server.
[0725] Step 3:
[0726] The server analyzes the received data and extracts important information. The input is a data stream sent from the terminal, and it analyzes voice tone and facial expression patterns using a generative AI model and machine learning algorithms. Specifically, the server runs an emotion analysis model using TensorFlow and analyzes this data. The output is structured data indicating the user's emotional state.
[0727] Step 4:
[0728] The server generates appropriate support instructions based on the analysis results. The input is structured data indicating the user's emotional state, and it uses generative AI to generate support messages based on prompts. Specifically, the server generates instructions such as "Observe the current situation calmly and take appropriate action" through a generative model. The output is optimized support instructions.
[0729] Step 5:
[0730] The server sends the generated support instructions to the terminal. The input is the support instructions generated by the server, which are sent to the terminal via the data transmission path. Specifically, the server repackets the data and sends it using a secure protocol. The output is the support instruction data received by the terminal.
[0731] Step 6:
[0732] The terminal displays and audibly presents received support instructions to the user. The input is data of support instructions sent from the server, which is then communicated to the user via the display and speaker. Specifically, the terminal displays messages on the screen and plays instructions using speech synthesis technology. The output is information conveyed to the user visually and aurally.
[0733] (Application Example 2)
[0734] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0735] In today's work environment, it is crucial to perform tasks efficiently while maintaining the health and safety of workers. However, conventional systems have struggled to monitor and respond to stress and fatigue levels in real time. Therefore, new methods are needed to ensure worker productivity and safety.
[0736] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0737] In this invention, the server includes means for collecting environmental information of the work area, means for analyzing worker emotional data in real time and generating feedback, and means for generating and notifying instructions suggesting work breaks according to the emotional state. This makes it possible to create an environment that provides optimal work instructions while taking into account the worker's health condition.
[0738] "Voice data" refers to electromagnetic wave data obtained digitally from the worker's voice and used for analysis.
[0739] "Video data" refers to visual information obtained as digital images of a worker's facial expressions and movements, which is then used for analysis.
[0740] A "terminal device" is a device used to acquire audio and video data and transmit it to the main control unit.
[0741] An "information transmission network" is a communication infrastructure used to transmit data acquired by terminal devices to a main control unit.
[0742] The "main control unit" is a central processing unit that analyzes collected data, extracts knowledge, and generates instructions.
[0743] "Knowledge" refers to information about actions and responses that should be taken, extracted as a result of data analysis.
[0744] "Feedback" refers to advice and notifications provided to workers based on analyzed emotional data.
[0745] "Stress level" is a quantitative indicator of a worker's mental state, showing how it affects their normal work performance.
[0746] A "suggestion to pause work" is an instruction that recommends a temporary break to alleviate the stress levels of workers.
[0747] The system implementing this invention mainly consists of terminal equipment and a main control unit (server). The terminal equipment is responsible for acquiring audio and video data and transmitting this data to the server via an information transmission network. The terminal is capable of monitoring the surrounding environment and emotional state of the user, who is a worker, in real time.
[0748] The server analyzes received audio and video data to extract knowledge about the worker's stress level and emotional state. This analysis uses machine learning models (e.g., TensorFlow) to apply generated intelligent processing algorithms, enabling automatic data recognition and action decisions. Based on these results, the server generates emotionally appropriate work instructions and feedback, which are then sent to the terminal device.
[0749] For example, if the server determines from the worker's facial expression data that their stress level is high, it can immediately send a message to the terminal suggesting a short break. This is an important means of supporting the mental health of workers while maintaining their work efficiency. Furthermore, the generative AI model used is constantly updated and designed to flexibly adapt to new situations.
[0750] An example of a prompt might be: "Recognize the worker's fatigue level, analyze the data, and generate appropriate feedback. Specifically, be able to suggest breaks when stress levels are high." Based on this prompt, the system will take action to provide the worker with the best possible support.
[0751] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0752] Step 1:
[0753] The terminal acquires the worker's voice and video data in real time. Here, the terminal's camera and microphone are used to record the worker's facial expressions and voice as digital data. The input is raw data from the camera and microphone, which is temporarily stored on the terminal.
[0754] Step 2:
[0755] The terminal transmits collected audio and video data to the server using an information transmission network. The input is the audio and video data stored on the terminal, which is converted and transmitted using a transfer protocol. The output is the data stream received on the server.
[0756] Step 3:
[0757] The server begins analyzing the received data. Here, a generative AI model is used to analyze the audio and video data and assess the worker's emotional state and stress level. The input is a stream of audio and video data, and the output is the assessed emotional state. The analysis involves machine learning algorithms, and the data is processed within the model.
[0758] Step 4:
[0759] The server generates feedback or instructions based on the analysis results. In this process, it recommends the most appropriate action to the worker based on pre-configured prompts. The input is an assessment of emotional state, and the output is the generated feedback or instruction message. Specifically, it might generate suggestions such as, "Why don't you take a short break?"
[0760] Step 5:
[0761] The server sends feedback or instructions to the terminal. Here, the generated message is displayed on the terminal's screen and also provided as audio guidance. The input is the generated instruction message, and the output is the visual and auditory information received by the user (worker).
[0762] Step 6:
[0763] Users adjust their actions based on the feedback they receive. This reduces worker stress and maintains an efficient work environment. Input is instructions from the terminal, and output is the user's actual actions. Specific actions include taking suggested breaks.
[0764] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0765] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0766] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0767] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0768] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0769] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0770] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0771] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0772] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0773] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0774] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0775] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0776] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0777] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0778] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0779] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0780] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0781] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0782] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0783] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0784] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0785] The following is further disclosed regarding the embodiments described above.
[0786] (Claim 1)
[0787] A means for collecting on-site environmental data using terminal devices that acquire acoustic and image signals,
[0788] Means for transmitting acquired acoustic and image signals to a central processing unit via a communication network,
[0789] In the central processing unit, means for analyzing the received signal and extracting important information,
[0790] A means for generating optimal rescue instructions based on extracted information and transmitting those instructions to a terminal device,
[0791] A means of displaying instructions transmitted from a terminal device to the user and supporting rescue operations,
[0792] A system that includes this.
[0793] (Claim 2)
[0794] The system according to claim 1, which applies a machine learning algorithm using generative artificial intelligence to analyze acoustic and image signals, and performs automatic situation recognition and action decision-making.
[0795] (Claim 3)
[0796] The system according to claim 1, wherein the central processing unit calculates the priority of rescue operations from the received acoustic and image signals and generates rescue instructions based on this.
[0797] "Example 1"
[0798] (Claim 1)
[0799] A means of collecting on-site environmental information using an information terminal that acquires sound and images,
[0800] Means for transmitting acquired sound and images to a central processing unit via a communication path,
[0801] In the central processing unit, means for analyzing received information and extracting important content,
[0802] A means for generating optimal rescue guidelines based on the extracted information and transmitting those guidelines to an information terminal,
[0803] A means of supporting relief efforts by presenting guidelines transmitted via information terminals to users,
[0804] A system that includes this.
[0805] (Claim 2)
[0806] The system according to claim 1, which applies machine learning methods using generative artificial intelligence to the analysis of sound and images to perform automatic situation recognition and action decisions.
[0807] (Claim 3)
[0808] The system according to claim 1, wherein a central processing unit calculates the priority of rescue operations from received sound and images, and generates rescue guidelines based on this.
[0809] "Application Example 1"
[0810] (Claim 1)
[0811] A means for collecting on-site environmental information using an information acquisition device that acquires acoustic and image signals,
[0812] Means for transmitting acquired acoustic and image signals to a data processing device via a communication network,
[0813] A data processing device includes means for analyzing received signals and extracting important information, including abnormal events.
[0814] A means for generating optimal action instructions based on extracted information and transmitting those instructions to an information acquisition device,
[0815] A means of displaying instructions transmitted by an information acquisition device to the user and supporting risk avoidance activities,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, which applies a learning algorithm using generative artificial intelligence to analyze acoustic and image signals, and performs environmental recognition and automatic action decision-making.
[0819] (Claim 3)
[0820] The system according to claim 1, wherein the data processing device calculates the priority of an action from the received acoustic signal and image signal, and generates an action instruction based on this.
[0821] "Example 2 of combining an emotion engine"
[0822] (Claim 1)
[0823] A means of collecting on-site situation information using an information terminal that acquires auditory and visual data,
[0824] Means for transmitting acquired auditory and visual data to a centralized information processing device via a data transmission path,
[0825] An information processing device includes means for analyzing received data and identifying and extracting important elements,
[0826] A means for generating optimal support instructions based on identified and extracted elements and transmitting those instructions to an information terminal,
[0827] A means of presenting instructions transmitted via an information terminal to the user and assisting in support tasks,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, which applies an automated learning algorithm using generative artificial intelligence to the analysis of auditory and visual data, and performs autonomous recognition of the situation and selection of actions.
[0831] (Claim 3)
[0832] The system according to claim 1, wherein the information processing device calculates the priority of support tasks from received auditory data and visual data, and generates support instructions based on this.
[0833] "Application example 2 when combining with an emotional engine"
[0834] (Claim 1)
[0835] A means for collecting environmental information of a work location using terminal equipment that acquires audio and video data,
[0836] Means for transmitting acquired audio and video data to a main control unit via an information transmission network,
[0837] The main control unit includes means for analyzing received information and extracting important knowledge,
[0838] A means for generating optimal work instructions based on extracted knowledge and transmitting those instructions to a terminal device,
[0839] A means of displaying instructions transmitted from a terminal device to the worker and supporting work activities,
[0840] A system that includes this.
[0841] (Claim 2)
[0842] The system according to claim 1, which applies machine learning using a generated intelligent processing algorithm to analyze audio data and video data to perform automatic situation recognition and action decision-making.
[0843] (Claim 3)
[0844] The system according to claim 1, wherein the main control unit calculates the priority of work activities from received audio data and video data, and generates work instructions based on this.
[0845] (Claim 4)
[0846] A means for analyzing workers' emotional data and generating real-time feedback corresponding to their emotional state,
[0847] A means of notifying workers of instructions, including suggestions for taking a break from work, based on stress levels estimated from emotional data,
[0848] The system according to claim 1, including the following: [Explanation of Symbols]
[0849] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for collecting on-site environmental data using terminal devices that acquire acoustic and image signals, Means for transmitting acquired acoustic and image signals to a central processing unit via a communication network, In the central processing unit, means for analyzing the received signal and extracting important information, A means for generating optimal rescue instructions based on extracted information and transmitting those instructions to a terminal device, A means of displaying instructions transmitted from a terminal device to the user and supporting rescue operations, A system that includes this.
2. The system according to claim 1, which applies a machine learning algorithm using generative artificial intelligence to analyze acoustic and image signals, thereby enabling automatic situation recognition and action decision-making.
3. The system according to claim 1, wherein the central processing unit calculates the priority of rescue operations from the received acoustic and image signals, and generates rescue instructions based on this.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A