system
The emergency notification system uses satellite communication and AI to convert voice to text, analyze images, and prioritize responses, addressing the challenges of identifying urgency and communication failures during disasters, ensuring rapid and accurate emergency assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional emergency reporting systems struggle during large-scale natural disasters or emergencies due to difficulties in identifying the position of the sender, judging urgency, and delayed responses caused by communication infrastructure failures.
An emergency notification system utilizing satellite communication to receive emergency calls, convert voice data to text using generation AI, analyze image data, and prioritize responses based on urgency, with real-time feedback for system optimization.
Enables rapid and accurate emergency response by determining urgency and coordinating rescue operations effectively, minimizing communication disruptions and improving response accuracy through continuous system adjustments.
Smart Images

Figure 2026070256000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present invention is to provide a new emergency reporting system for preventing victims from being isolated in situations where ground communication infrastructure fails during large-scale natural disasters or emergencies. In conventional systems, there are problems that it is difficult to identify the position of the sender and judge the urgency, and furthermore, the response is delayed due to the concentration of reports during large-scale disasters.
Means for Solving the Problems
[0005] This invention provides a means for receiving emergency calls using satellite communication, converting voice data into text using generation AI, and analyzing image data to quickly and accurately grasp the situation of disaster victims. This enables the aggregation of analyzed data, prioritization based on urgency, and prompt notification to emergency services. Furthermore, it proposes a means to improve the accuracy and efficiency of the response by using voice recognition technology to determine the urgency, re-analyzing as necessary, and adjusting the system based on feedback from emergency services.
[0006] "Satellite communication" is a method of communication that transmits and receives information via artificial satellites, without using terrestrial communication infrastructure.
[0007] An "emergency call" is the transmission of information to quickly request rescue in the event of an emergency such as a disaster or accident.
[0008] "Generative AI" is an artificial intelligence technology that uses machine learning algorithms to analyze data and generate new information.
[0009] "Speech recognition technology" is a technology that analyzes speech data and converts it into text information.
[0010] An "image analysis algorithm" is a computational method for processing image data and identifying specific patterns or features.
[0011] Prioritization is the act of determining the order in which to process a large number of tasks or pieces of information.
[0012] "Notification means" refers to methods or devices for transmitting specific information to a target person.
[0013] "Feedback" is the process of reacting to received information and making responses or adjustments. [Brief explanation of the drawing]
[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention relates to an embodiment of an emergency notification system utilizing satellite communication. When a user sends an SOS message using their smartphone during a disaster or emergency, the system receives the data via satellite communication. The received data is then processed by a server.
[0036] The server first uses a generation AI to convert audio data into text. Machine learning techniques are used to extract keywords indicating urgency from the audio and record them as text. Simultaneously, image analysis algorithms are used to analyze the situation of disaster victims and changes in the environment from photos and videos. For example, if a user's image shows collapsed buildings or injured people, the server identifies them and calculates the urgency as further information.
[0037] The server then aggregates the analyzed data and determines a priority for each emergency call. It selects the most important calls from a large number and sorts them in order of the need for immediate response. This process allows for an efficient response to situations requiring emergency attention. For example, calls from seriously injured individuals or high-risk areas are prioritized within the system, and rescue operations are coordinated to ensure rapid response.
[0038] Prioritized information is sent from the server to emergency services. The server then refers to a database of pre-registered emergency services and sends the necessary information to the most appropriate service provider. This notification includes the user's location, the extent of the damage, and recommended emergency responses, enabling the emergency services to make informed decisions.
[0039] As the final step in this embodiment, the server optimizes the system based on feedback received after the notification. Based on the feedback provided by the emergency services, the system can make real-time adjustments to further improve accuracy. This improves the speed and accuracy of responses in similar cases.
[0040] Because this system does not rely on conventional terrestrial communications, it minimizes the impact of communication disruptions during disasters and enables the rapid rescue of victims.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] When a user creates an emergency call or SOS message using a dedicated application on their smartphone and presses the send button, the message is sent. At this point, the user's location information and message content are transmitted from the device via satellite communication.
[0044] Step 2:
[0045] The server acquires emergency call data received via satellite communication. Immediately after reception, the data is transferred to an analysis module. Here, the audio and image data are ready to be processed separately.
[0046] Step 3:
[0047] The server uses a generation AI to convert voice data into text. Using speech recognition technology, it extracts keywords and phrases indicating urgency, which serve as the basis for determining the urgency of the report. Simultaneously, it uses an analysis algorithm to analyze the environment and injury status from image data, obtaining information necessary for accurate situation assessment.
[0048] Step 4:
[0049] The server centralizes the analyzed information and prioritizes each report. An AI algorithm automatically prioritizes rescue cases based on the urgency and location of the victims. Furthermore, detailed disaster patterns based on the user's equipment and surrounding environment are simultaneously analyzed.
[0050] Step 5:
[0051] The server promptly notifies emergency services. The notification includes special details about the disaster, location information, and urgency ranking, providing data that enables immediate response. The notification protocol is automatically configured based on a pre-registered database of emergency services.
[0052] Step 6:
[0053] The server continuously monitors and gathers feedback after an alert is received. Based on feedback from emergency services, the system is continuously adjusted, and AI algorithms are improved and databases are updated in real time. Through this process, the overall accuracy and efficiency of the system's response are enhanced.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] During disasters and emergencies, there is a need to ensure that victims can receive emergency assistance quickly and effectively, even when ground-based communication infrastructure is unavailable. Conventional systems have the problem of hindering the transmission of emergency calls due to communication disruptions. Furthermore, there is a challenge in that there are insufficient means to accurately assess the urgency of received information and quickly coordinate with the appropriate emergency services, leading to delays in response.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes a device that receives emergency calls using satellite communication, a device that converts voice information into text and analyzes image information using generative artificial intelligence, and a device that aggregates the analyzed information and determines the processing order based on the urgency. This makes it possible to quickly determine the urgency and transmit information to the appropriate emergency agency even if ground communication is interrupted.
[0059] "Satellite communication" is a technology that transmits and receives information via communication satellites without using terrestrial communication infrastructure.
[0060] An "emergency call" is communication information sent to request help in emergency situations such as disasters or getting lost.
[0061] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to analyze data and generate information such as speech and text.
[0062] "Converting audio information to text" is the process of analyzing audio and representing it as a corresponding string of characters.
[0063] "Analyzing image information" is a technique that uses image data to identify situations and objects and extract meaningful information.
[0064] "Device" refers to hardware or software designed to perform a series of processes.
[0065] "Urgency" is an indicator that shows the degree to which an emergency response is necessary, and it is a criterion for determining the order in which a rapid response is required.
[0066] "Processing order" refers to guidelines for determining the order in which multiple tasks are performed based on their priority.
[0067] An "emergency agency" is a public or private organization that has the facilities and personnel to respond to emergencies such as disasters and accidents.
[0068] This invention is an emergency notification system that utilizes satellite communication to enable users to receive rapid assistance in emergency situations.
[0069] First, the user uses a dedicated smartphone app. This app is used to send SOS messages in the event of a disaster or emergency. The device is equipped with a satellite communication module and transmits signals via an airborne communication network without using ground-based communication infrastructure. A general-purpose module that supports various protocols and standards could be used as the satellite communication module.
[0070] The server processes data received via satellite communication. Generative artificial intelligence is used to transcribe audio data into text. A general-purpose AI model that excels at speech recognition and can perform effective text conversion is considered. Next, machine learning algorithms are applied to the text data to extract specific keywords indicating urgency. For image data, commonly used image analysis algorithms are used to determine the disaster situation from the video.
[0071] The analyzed data is aggregated as an urgency score, and the processing order is determined based on its importance. This information is notified to the nearest emergency services, and instructions are given to begin immediate action. This process enables rapid and accurate emergency assistance.
[0072] For example, if a user makes a voice call saying, "My house is on fire, please help," the voice data is converted to text by the server, and keywords such as "fire" and "help" are extracted as factors that increase the urgency. As a result, the server processes it as a high-priority call and can immediately transmit the information to fire departments and other relevant organizations.
[0073] Example of a prompt:
[0074] "A user has sent an SOS message from a disaster area. Analyze the audio and video, calculate the urgency, and notify the appropriate agency."
[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0076] Step 1:
[0077] Procedure for users to send SOS messages
[0078] Users launch a dedicated app on their smartphones to record newly occurring emergencies. A voice input function allows them to enter comments describing the situation using voice. They can also capture images and videos of the scene using the camera function and attach them to the SOS message. This data is transmitted from the device to the server via a satellite communication module. Input consists of voice and image data, while output is structured data sent to the server.
[0079] Step 2:
[0080] The server receives the data and converts the audio data into text.
[0081] The server receives the SOS message sent from the terminal. The received audio data is converted into text using a generative AI model. Specifically, the audio waveform data is input to the AI model, which is then analyzed and output as text data.
[0082] Step 3:
[0083] The server extracts urgent keywords from the text.
[0084] The server analyzes the text data and extracts keywords indicating urgency (e.g., "Help," "Fire"). It applies a machine learning algorithm to identify important keywords from a pre-trained model. The input is text data, and the output is the extracted urgency keywords.
[0085] Step 4:
[0086] The process by which the server analyzes image data.
[0087] The server analyzes the received image data to determine the situation and environment of the subject. Using image analysis algorithms, it identifies, for example, collapsed buildings or the condition of people. The input is image data, and the output is the analyzed situational information.
[0088] Step 5:
[0089] A procedure in which the server calculates urgency and determines the processing order.
[0090] The server integrates keywords from voice analysis with image analysis results to calculate the overall urgency. It generates an urgency score using a proprietary calculation model and compares it to other reports based on that score. The output is this score, which is used to determine the processing order.
[0091] Step 6:
[0092] The server notifies the most appropriate emergency agency.
[0093] Based on the determined processing order, the server notifies the most appropriate emergency agency. The information sent to each emergency agency includes the user's location information, transcribed audio information, and analyzed image information. This allows the receiving agency to quickly understand the situation and begin responding. The output is notification data.
[0094] (Application Example 1)
[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0096] In natural disasters and emergencies, there is a need for communication methods that enable rapid and accurate reporting and initial response. However, conventional systems that rely on terrestrial communications struggle to effectively report emergencies when communication infrastructure is destroyed. Furthermore, they lack intuitive operation methods to enable users to report emergencies at the appropriate time. In addition, there are few means to effectively determine the priority of emergency calls, hindering efficient rescue operations.
[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0098] In this invention, the server includes a receiving means for receiving emergency calls using satellite communication, an analysis means for converting voice data into text using a generation AI and analyzing image data, and a prioritizing means for aggregating the analyzed data and determining priorities based on urgency. This realizes an emergency call system that does not depend on communication infrastructure, allows users to make calls intuitively, and enables rapid rescue operations through efficient prioritization.
[0099] "Satellite communication" is a technology that transmits and receives communication signals via artificial satellites, making it possible to transmit signals over a wide area without relying on ground-based communication infrastructure.
[0100] "Generative AI" is an artificial intelligence technology that automatically generates information from data, and has the ability to analyze various types of data such as audio and images to extract text and features.
[0101] "Converting audio data to text" is the process of analyzing audio information and converting it into corresponding textual information.
[0102] "Analyzing image data" refers to the technique of extracting information from images and identifying specific features or patterns.
[0103] "Determining priorities" means evaluating the importance of various data and reports and setting the order in which they require immediate attention.
[0104] "Recording means installed in a portable display device" refers to a function incorporated into a display device that can be easily carried by an individual, which acquires and records audio and video.
[0105] A "startup mechanism" is a function that automatically starts the system when a user performs an action that meets specific conditions.
[0106] A "notification mechanism" is a function for appropriately transmitting analyzed information to relevant authorities and agencies.
[0107] To realize this invention, a portable display device carried by the user, a server, and a satellite communication network are required. The portable display device includes a microphone and camera necessary for acquiring audio and video. When the user gives a voice command in an emergency, the device automatically starts recording audio and video data. This data is transmitted from the device to the server via satellite communication.
[0108] The server uses a generative AI model to convert audio data into text and extracts keywords from that text to assess urgency. Specifically, it uses Google® Speech-to-Text to convert audio to text and then analyzes the extracted text with TENSORFLOW®. In addition, it uses image processing technologies such as OpenCV to evaluate environmental changes and the degree of danger in video data.
[0109] Based on the analyzed data, the server prioritizes emergency events and generates a list of notifications, ordered from those requiring immediate attention to those requiring it. This enables efficient rescue operations. By utilizing satellite communication, rather than relying on existing communication infrastructure, a highly reliable notification system can be established even in the event of communication disruptions.
[0110] As a concrete example, in a mountainous area where ground communication is cut off due to a natural disaster, if a hiker is about to fall, they can simply say "SOS" into a portable display device, and the appropriate agency will be quickly notified via satellite communication.
[0111] An example of a prompt message for a generative AI model to analyze audio data is: "Design a machine learning model to analyze audio data and extract urgency keywords. Also, propose a method for evaluating the level of risk through image analysis."
[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0113] Step 1:
[0114] The user speaks "SOS" into the portable display device to report an emergency. The device records this audio as sound data and simultaneously acquires video data using its camera. The audio and video inputs are used to make the emergency call.
[0115] Step 2:
[0116] The terminal transmits acquired audio and video data to the server via satellite communication. The input is data recorded within the terminal, which becomes the information to be transmitted to the server. The server receives this data and prepares for data analysis.
[0117] Step 3:
[0118] The server uses a generative AI model to convert speech data into text. It uses Google Speech-to-Text to generate corresponding strings from the speech data and extract urgent keywords. The input is speech data, and the output is textual information. In this process, the user's spoken content is converted into specific textual information.
[0119] Step 4:
[0120] The server uses image analysis techniques such as OpenCV to analyze video data and evaluate the environmental conditions and changes. The input is video data, and the output is feature information of the analyzed video. In this step, high-risk elements within the video are identified.
[0121] Step 5:
[0122] The server integrates the results of voice-to-text and video analysis to determine the priority of notifications based on urgency. This process prioritizes notifications by considering emergency keywords and environmental analysis information. The input is transcribed voice information and video feature information, and the output is prioritized notification data.
[0123] Step 6:
[0124] The server notifies the appropriate emergency services of prioritized alert data. The input is prioritized alert information, which becomes the content of the notification to the emergency services. The server refers to the database, selects the most suitable service based on the content of each piece of information, and sends a notification to encourage a rapid response.
[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0126] This invention relates to an embodiment of an emergency notification system for disasters and emergencies that incorporates an emotion engine to recognize the user's emotions. In this embodiment, the user inputs an emergency message using a dedicated application on their smartphone. The terminal collects emotion data along with voice and image data and transmits it to a server via satellite communication.
[0127] The transmitted data is processed on the server using an analysis method that utilizes generation AI. Within this process, an emotion engine recognizes the user's emotions through voice analysis, and this information is used to determine the level of urgency. For example, based on changes in the user's voice tone and volume, features indicating heightened emotions are extracted and considered as factors that increase the level of urgency.
[0128] Furthermore, image analysis is performed to determine the environment and the extent of injuries from visual information. This information is integrated on a server, and a prioritization mechanism determines the priority of the response. The emotion engine evaluates the impact of emotional information on the severity and urgency of the report by comparing it with past databases.
[0129] Next, the server notifies emergency services of the prioritized information. This process also includes emotional data, enabling emergency services to respond appropriately, taking into account the psychological state of the victims. For example, based on emotional information, emergency services can instruct immediate action on cases that require priority.
[0130] Ultimately, a feedback mechanism improves the accuracy of the operational system through feedback from emergency services. Based on this feedback, the emotion engine algorithm is improved, and the system is continuously adjusted to achieve even greater accuracy in future incidents. This results in a comprehensive emergency call system that enables rapid and highly accurate responses.
[0131] The following describes the processing flow.
[0132] Step 1:
[0133] The user enters an emergency call message and records a voice message through a smartphone application. The device inputs the user's voice into an emotion engine in real time and acquires data to estimate their emotions.
[0134] Step 2:
[0135] The device packages the acquired voice messages, emotion data, and location information and transmits them to the server via satellite communication. The transmitted information also includes the results of the user's emotion analysis.
[0136] Step 3:
[0137] The server inputs the received data into an analysis module and uses a generation AI to convert the voice message into text. Furthermore, it evaluates the user's psychological urgency based on the emotional data provided by the emotion engine.
[0138] Step 4:
[0139] The server analyzes the collected image data using image analysis algorithms to extract information about the damage and the environment. It then references recognized emotion data to comprehensively assess the overall urgency of the situation.
[0140] Step 5:
[0141] The server centralizes the analyzed urgency and sentiment data and uses prioritization mechanisms to determine the priority of rescue operations. Based on this information, cases requiring a rapid response are identified.
[0142] Step 6:
[0143] The server notifies emergency services of priority information. This notification includes detailed disaster information as well as information about the user's emotional state, allowing emergency services to accurately understand the situation and take appropriate action when psychological support is needed.
[0144] Step 7:
[0145] The server continuously monitors feedback from emergency calls and adjusts its emotion engine and analysis methods based on reports from emergency services. This improves the system so that it can provide a more accurate emergency response in the future.
[0146] (Example 2)
[0147] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0148] In the event of a disaster or emergency, accurate analysis of on-site reports is essential for a swift and appropriate response. However, conventional systems lack the accuracy to analyze voice and image data, making it particularly difficult to determine the level of urgency while considering the user's emotional state. Furthermore, coordinating with emergency services presents challenges in responding to the specific needs of the field.
[0149] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0150] In this invention, the server includes a communication device that receives emergency calls using satellite communication, an analysis mechanism that converts voice signals into text data and analyzes image information using generation AI technology, and an emergency assessment device that recognizes the user's emotions and evaluates the urgency by integrating that emotion information with the analyzed data. This makes it possible to accurately determine the priority of responses in the event of a disaster or distress and to provide a swift and appropriate emergency response.
[0151] A "communication device" is a device that uses satellite communication to receive emergency calls from outside.
[0152] An "analysis mechanism" is a device that uses generative AI technology to convert audio signals into text data or to perform detailed analysis of image information.
[0153] An "urgency assessment device" is a device that analyzes the user's emotions, evaluates the local situation along with the integrated data, and determines the level of urgency.
[0154] A "prioritization device" is a device that determines the priority of responses based on the urgency level obtained by an urgency assessment device.
[0155] A "communication device" is a device used to accurately and quickly notify relevant organizations of emergency information.
[0156] A "regulating device" is a device used to improve the analysis mechanism based on feedback from related organizations, thereby enhancing the overall accuracy of the system.
[0157] This invention is a system for effectively making emergency calls during disasters or emergencies. The system consists of a mechanism in which a user inputs an emergency message using a smartphone terminal, and this information is transmitted to a server via satellite communication.
[0158] The device is equipped with the functionality to collect voice and image data. Users use a dedicated smartphone application to make voice notifications and send images. The device utilizes speech recognition technology to convert the user's voice into text data. This conversion uses a generative AI model, providing more advanced speech analysis capabilities.
[0159] The server receives data transmitted from the terminal and processes it using an analysis mechanism. This analysis mechanism, which includes a generative AI, not only converts audio data into text but also has the ability to identify emotions from factors such as tone and volume. Furthermore, it can evaluate the user's visual situation using image data.
[0160] For example, consider a scenario where a user is lost in a mountainous area and uses their device to record a voice message saying, "Please help me, I don't know where I am," and sends it along with an image. An example of a prompt message in this case would be, "We have received your emergency call, will analyze the user's emotions from the voice data, and will develop a rapid response plan."
[0161] The server's urgency assessment device quickly determines the urgency of the situation based on analyzed emotional and visual information. Based on this, the priority determination device organizes the response priorities and notifies those with the highest level of urgency first.
[0162] Finally, the server notifies the relevant organizations of appropriate response information through a communication device. The communication device conveys detailed information about the report, including the user's emotional state, to the relevant organizations, helping them to implement the most appropriate response on the scene.
[0163] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0164] Step 1:
[0165] The user enters an emergency message using a dedicated smartphone application and records audio and image data to the device. This input includes the user's voice and video captured by the smartphone camera. The device formats this data and saves it in an appropriate format for subsequent processing.
[0166] Step 2:
[0167] The device processes recorded audio data and converts it into text data using a generation AI model. The input is the user's audio file, which is then passed through an analysis engine to convert it into text information. This outputs specific textual information. The intonation and tone of the voice are also analyzed simultaneously, and characteristic data indicating emotion is extracted.
[0168] Step 3:
[0169] The device analyzes stored image data and extracts visual features. The input is an image file taken by the user, and an image processing algorithm is used to evaluate the environmental conditions and the state of people. The output includes visual hazards and environmental information derived from the image.
[0170] Step 4:
[0171] The terminal aggregates the converted text data, sentiment data, and image analysis results and sends them to the server via the internet. This transmission process involves data integration and packetization to ensure secure and rapid transmission to the server.
[0172] Step 5:
[0173] The server processes the received data using an analysis mechanism. It analyzes text data to identify urgent elements and evaluates the user's psychological state based on emotional data extracted from the audio. This process efficiently identifies information with high urgency.
[0174] Step 6:
[0175] The server integrates the emotional and image information analyzed by the urgency assessment device to calculate an urgency score. In this step, the urgency of each input data is quantified and prioritized. The final output is a response plan with determined priorities.
[0176] Step 7:
[0177] The server notifies relevant organizations of prioritized information via communication devices. This notification includes the user's location and instructions for responding to high-priority cases. This allows for immediate and appropriate responses on the ground.
[0178] (Application Example 2)
[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0180] Conventional emergency reporting systems often judge the urgency of a situation based on a uniform standard without considering the caller's psychological state or emotions, which can lead to a lack of prompt and appropriate responses tailored to the actual situation. As a result, there is a problem where the prioritization of emergency responses is incorrect, and immediate responses may be delayed. Furthermore, responding without considering the emotional state of the victims can increase their psychological burden.
[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0182] In this invention, the server includes means for receiving emergency calls using satellite communication, means for converting voice information into text using a generation AI and analyzing video information, means for aggregating the analyzed information and determining a ranking based on the degree of urgency, and means for analyzing the user's emotions using an emotion engine and considering this in determining the degree of urgency. This makes it possible to evaluate the overall degree of urgency, including the emotional state of the caller, enabling faster and more appropriate prioritization and response.
[0183] "Satellite communication" is a means of communication that uses artificial satellites to send and receive information between devices located in distant places on Earth.
[0184] An "emergency call receiving device" is a device that quickly receives emergency messages from callers in the event of an abnormal situation or danger.
[0185] "Generative AI" refers to artificial intelligence technology that can analyze various types of data, including audio and video, and extract useful information from them.
[0186] "Auditory information" refers to data that records the words and sounds spoken by a person, and it serves as the basis for analyzing context and emotions.
[0187] "Visual information" refers to visual data acquired through cameras and other imaging devices, and is information that visually captures the state of objects and the environment.
[0188] "Analysis methods" refer to technical techniques for processing collected data and extracting meaningful information from it.
[0189] A "system for determining priority based on urgency" is a system for evaluating the seriousness and urgency of a situation and determining the priority of responses accordingly.
[0190] An "emotion engine" is a technological element that can identify and analyze a person's emotional state from audio and video.
[0191] "The caller's emotional state" refers to the feelings and psychological condition of the person making the emergency call.
[0192] "User" refers to the individual or organization that uses this emergency call system.
[0193] The system that realizes this invention consists of the cooperation of multiple devices and software. When making an emergency call, the user uses a portable information terminal. This terminal is equipped with satellite communication capabilities, enabling highly reliable communication unaffected by the external environment.
[0194] When a user presses the emergency button, the device's microphone records audio and the camera captures video. This information is processed in real time by a generative AI; the audio is converted to text and the video is analyzed. The generative AI uses the Google Cloud Speech-to-Text API to transcribe the audio, while the Google Cloud Vision API is used to analyze the video. This enables detailed analysis of both audio and video.
[0195] The server receives information transmitted from the terminal and analyzes the user's emotions using an emotion engine. In this process, it detects changes in the tone and volume of the user's voice to determine the level of emotional arousal. It also compares this data with past data to determine the degree of urgency. The server sets priorities according to the degree of urgency and immediately notifies emergency services if necessary.
[0196] For example, if a user witnesses someone behaving suspiciously on the street, they can report the emergency situation through the application. Information reflecting the urgency of the voice report is immediately sent to the server, and if action is needed, instructions are quickly given to emergency services. Through this process, it is possible to ensure an appropriate and rapid response in emergencies.
[0197] An example of a prompt to smoothly initiate the analysis of a generative AI model would be: "Based on recent emergency call cases, analyze the emotions in this voice data. Provide guidance to assess the urgency and select the appropriate response route." This allows for more accurate assessment of emotional changes and urgency.
[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0199] Step 1:
[0200] When a user presses the emergency call button on their mobile device, the device begins recording audio and capturing video with its camera. The input consists of the user's audio and video data, which is then converted into a format that can be processed by the AI model. Specifically, this involves collecting the necessary data using the device's microphone and camera.
[0201] Step 2:
[0202] The device uses generative AI to convert recorded audio information into text. This involves analyzing the audio data using the Google Cloud Speech-to-Text API and converting it into text. The input is audio data, and the output is text data. The process includes extracting audio features within the device and processing them as text strings.
[0203] Step 3:
[0204] The device analyzes captured video information using the Google Cloud Vision API. The input is video data, and the output is information derived from the analysis of various visual elements. Specifically, it recognizes objects and environmental characteristics within the image and collects this information as data.
[0205] Step 4:
[0206] The terminal transmits the acquired text information and video analysis results to the server. This is done using satellite communication, ensuring highly reliable data transfer. The input is the transcribed audio information and analysis results, and the output is the information sent to the server.
[0207] Step 5:
[0208] The server analyzes the user's emotions using an emotion engine based on the information it receives. For example, it can infer stress and tension from factors such as voice tone and volume. The input is the user's voice tone information, and the output is the analyzed emotion information. The generation of emotion data within the server is a specific operation.
[0209] Step 6:
[0210] The server aggregates the analysis results and sets priorities based on urgency. During this process, the analysis results are compared with other emergency call data to determine which response should be given the highest priority. The input consists of analyzed emotional information and video data, and the output is the priority ranking. The main operation involves data comparison processing performed internally by the server.
[0211] Step 7:
[0212] Based on the priority determined by the server, a notification is sent to emergency services. This notification includes audio data, video data, and an analyzed level of urgency, enabling an appropriate response. The input is the data of the set priority, and the output is the notification to emergency services. The automatic transmission of data from the server is the specific operation in this processing step.
[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0216] [Second Embodiment]
[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0229] This invention relates to an embodiment of an emergency notification system utilizing satellite communication. When a user sends an SOS message using their smartphone during a disaster or emergency, the system receives the data via satellite communication. The received data is then processed by a server.
[0230] The server first uses a generation AI to convert audio data into text. Machine learning techniques are used to extract keywords indicating urgency from the audio and record them as text. Simultaneously, image analysis algorithms are used to analyze the situation of disaster victims and changes in the environment from photos and videos. For example, if a user's image shows collapsed buildings or injured people, the server identifies them and calculates the urgency as further information.
[0231] The server then aggregates the analyzed data and determines a priority for each emergency call. It selects the most important calls from a large number and sorts them in order of the need for immediate response. This process allows for an efficient response to situations requiring emergency attention. For example, calls from seriously injured individuals or high-risk areas are prioritized within the system, and rescue operations are coordinated to ensure rapid response.
[0232] Prioritized information is sent from the server to emergency services. The server then refers to a database of pre-registered emergency services and sends the necessary information to the most appropriate service provider. This notification includes the user's location, the extent of the damage, and recommended emergency responses, enabling the emergency services to make informed decisions.
[0233] As the final step in this embodiment, the server optimizes the system based on feedback received after the notification. Based on the feedback provided by the emergency services, the system can make real-time adjustments to further improve accuracy. This improves the speed and accuracy of responses in similar cases.
[0234] Because this system does not rely on conventional terrestrial communications, it minimizes the impact of communication disruptions during disasters and enables the rapid rescue of victims.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] When a user creates an emergency call or SOS message using a dedicated application on their smartphone and presses the send button, the message is sent. At this point, the user's location information and message content are transmitted from the device via satellite communication.
[0238] Step 2:
[0239] The server acquires emergency call data received via satellite communication. Immediately after reception, the data is transferred to an analysis module. Here, the audio and image data are ready to be processed separately.
[0240] Step 3:
[0241] The server uses a generation AI to convert voice data into text. Using speech recognition technology, it extracts keywords and phrases indicating urgency, which serve as the basis for determining the urgency of the report. Simultaneously, it uses an analysis algorithm to analyze the environment and injury status from image data, obtaining information necessary for accurate situation assessment.
[0242] Step 4:
[0243] The server centralizes the analyzed information and prioritizes each report. An AI algorithm automatically prioritizes rescue cases based on the urgency and location of the victims. Furthermore, detailed disaster patterns based on the user's equipment and surrounding environment are simultaneously analyzed.
[0244] Step 5:
[0245] The server promptly notifies emergency services. The notification includes special details about the disaster, location information, and urgency ranking, providing data that enables immediate response. The notification protocol is automatically configured based on a pre-registered database of emergency services.
[0246] Step 6:
[0247] The server continuously monitors and gathers feedback after an alert is received. Based on feedback from emergency services, the system is continuously adjusted, and AI algorithms are improved and databases are updated in real time. Through this process, the overall accuracy and efficiency of the system's response are enhanced.
[0248] (Example 1)
[0249] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0250] During disasters and emergencies, there is a need to ensure that victims can receive emergency assistance quickly and effectively, even when ground-based communication infrastructure is unavailable. Conventional systems have the problem of hindering the transmission of emergency calls due to communication disruptions. Furthermore, there is a challenge in that there are insufficient means to accurately assess the urgency of received information and quickly coordinate with the appropriate emergency services, leading to delays in response.
[0251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0252] In this invention, the server includes a device that receives emergency calls using satellite communication, a device that converts voice information into text and analyzes image information using generative artificial intelligence, and a device that aggregates the analyzed information and determines the processing order based on the urgency. This makes it possible to quickly determine the urgency and transmit information to the appropriate emergency agency even if ground communication is interrupted.
[0253] "Satellite communication" is a technology that transmits and receives information via communication satellites without using terrestrial communication infrastructure.
[0254] An "emergency call" is communication information sent to request help in emergency situations such as disasters or getting lost.
[0255] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to analyze data and generate information such as speech and text.
[0256] "Converting audio information to text" is the process of analyzing audio and representing it as a corresponding string of characters.
[0257] "Analyzing image information" is a technique that uses image data to identify situations and objects and extract meaningful information.
[0258] "Device" refers to hardware or software designed to perform a series of processes.
[0259] "Urgency" is an indicator that shows the degree to which an emergency response is necessary, and it is a criterion for determining the order in which a rapid response is required.
[0260] "Processing order" refers to guidelines for determining the order in which multiple tasks are performed based on their priority.
[0261] An "emergency agency" is a public or private organization that has the facilities and personnel to respond to emergencies such as disasters and accidents.
[0262] This invention is an emergency notification system that utilizes satellite communication to enable users to receive rapid assistance in emergency situations.
[0263] First, the user uses a dedicated smartphone app. This app is used to send SOS messages in the event of a disaster or emergency. The device is equipped with a satellite communication module and transmits signals via an airborne communication network without using ground-based communication infrastructure. A general-purpose module that supports various protocols and standards could be used as the satellite communication module.
[0264] The server processes data received via satellite communication. Generative artificial intelligence is used to transcribe audio data into text. A general-purpose AI model that excels at speech recognition and can perform effective text conversion is considered. Next, machine learning algorithms are applied to the text data to extract specific keywords indicating urgency. For image data, commonly used image analysis algorithms are used to determine the disaster situation from the video.
[0265] The analyzed data is aggregated as an urgency score, and the processing order is determined based on its importance. This information is notified to the nearest emergency services, and instructions are given to begin immediate action. This process enables rapid and accurate emergency assistance.
[0266] For example, if a user makes a voice call saying, "My house is on fire, please help," the voice data is converted to text by the server, and keywords such as "fire" and "help" are extracted as factors that increase the urgency. As a result, the server processes it as a high-priority call and can immediately transmit the information to fire departments and other relevant organizations.
[0267] Example of a prompt:
[0268] "A user has sent an SOS message from a disaster area. Analyze the audio and video, calculate the urgency, and notify the appropriate agency."
[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0270] Step 1:
[0271] Procedure for users to send SOS messages
[0272] Users launch a dedicated app on their smartphones to record newly occurring emergencies. A voice input function allows them to enter comments describing the situation using voice. They can also capture images and videos of the scene using the camera function and attach them to the SOS message. This data is transmitted from the device to the server via a satellite communication module. Input consists of voice and image data, while output is structured data sent to the server.
[0273] Step 2:
[0274] The server receives the data and converts the audio data into text.
[0275] The server receives the SOS message sent from the terminal. The received audio data is converted into text using a generative AI model. Specifically, the audio waveform data is input to the AI model, which is then analyzed and output as text data.
[0276] Step 3:
[0277] The server extracts urgent keywords from the text.
[0278] The server analyzes the text data and extracts keywords indicating urgency (e.g., "Help," "Fire"). It applies a machine learning algorithm to identify important keywords from a pre-trained model. The input is text data, and the output is the extracted urgency keywords.
[0279] Step 4:
[0280] The process by which the server analyzes image data.
[0281] The server analyzes the received image data to determine the situation and environment of the subject. Using image analysis algorithms, it identifies, for example, collapsed buildings or the condition of people. The input is image data, and the output is the analyzed situational information.
[0282] Step 5:
[0283] Procedure for the server to calculate the urgency and determine the processing order
[0284] The server integrates the keywords from voice analysis and the image analysis results to calculate the overall urgency. It generates an urgency score using its own calculation model and compares it with other reports based on this score. The output is this score, which is used to set the processing order.
[0285] Step 6:
[0286] Procedure for the server to notify the optimal emergency agency
[0287] Based on the determined processing order, the server notifies the optimal emergency agency. The information sent to each emergency agency includes the location information sent by the user, the texturized voice information, and the analyzed image information. This enables the receiving agency to quickly understand the situation and start responding. The output is the notification data.
[0288] (Application Example ¹)
[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0290] In natural disasters and emergencies, there is a need for communication means that enable quick and accurate reporting and initial response. However, in conventional systems that rely on terrestrial communication, effective reporting is difficult when the communication infrastructure is damaged. Also, there is a lack of intuitive operation methods for users to report at an appropriate timing. Furthermore, there are few means to effectively determine the priority order of emergency reports, which hinders efficient rescue activities.
[0291] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.
[0292] In this invention, the server includes a receiving means for receiving emergency calls using satellite communication, an analysis means for converting voice data into text using a generation AI and analyzing image data, and a prioritizing means for aggregating the analyzed data and determining priorities based on urgency. This realizes an emergency call system that does not depend on communication infrastructure, allows users to make calls intuitively, and enables rapid rescue operations through efficient prioritization.
[0293] "Satellite communication" is a technology that transmits and receives communication signals via artificial satellites, making it possible to transmit signals over a wide area without relying on ground-based communication infrastructure.
[0294] "Generative AI" is an artificial intelligence technology that automatically generates information from data, and has the ability to analyze various types of data such as audio and images to extract text and features.
[0295] "Converting audio data to text" is the process of analyzing audio information and converting it into corresponding textual information.
[0296] "Analyzing image data" refers to the technique of extracting information from images and identifying specific features or patterns.
[0297] "Determining priorities" means evaluating the importance of various data and reports and setting the order in which they require immediate attention.
[0298] "Recording means installed in a portable display device" refers to a function incorporated into a display device that can be easily carried by an individual, which acquires and records audio and video.
[0299] A "startup mechanism" is a function that automatically starts the system when a user performs an action that meets specific conditions.
[0300] A "notification mechanism" is a function for appropriately transmitting analyzed information to relevant authorities and agencies.
[0301] To implement this invention, a portable display device carried by a user, a server, and a satellite communication network are required. The portable display device includes a microphone and a camera necessary for acquiring sound and video. When the user instructs a voice report in an emergency, the terminal automatically starts recording voice and video data. These data are transmitted from the terminal to the server via satellite communication.
[0302] The server uses a generative AI model to convert voice data into text and extracts phrases for evaluating the urgency from the text. Specifically, Google Speech-to-Text is used to convert voice to text, and the extracted text is analyzed with TensorFlow. Also, video data is evaluated for environmental changes and danger levels using image processing techniques such as OpenCV.
[0303] Based on the analyzed data, the server determines the priority of the emergency event and generates a report list set in order from those that require immediate response. This enables efficient rescue activities. By utilizing satellite communication without relying on the communication infrastructure, a highly reliable reporting system can be constructed even during communication disruptions.
[0304] As a specific example, even when a hiker is about to slip in a mountainous area where ground communication means have been cut off due to a natural disaster, by uttering "SOS" towards the portable display device, a notification is quickly sent to the appropriate agency via satellite communication.
[0305] An example of a prompt sentence for analyzing voice data with a generative AI model is "Please design a machine learning model for analyzing voice data and extracting the included urgency keywords. Also, please propose a method for evaluating the danger level through image analysis."
[0306] The flow of specific processing in Application Example 1 will be described using FIG. 12.
[0307] Step 1:
[0308] The user speaks "SOS" into the portable display device to report an emergency. The device records this audio as sound data and simultaneously acquires video data using its camera. The audio and video inputs are used to make the emergency call.
[0309] Step 2:
[0310] The terminal transmits acquired audio and video data to the server via satellite communication. The input is data recorded within the terminal, which becomes the information to be transmitted to the server. The server receives this data and prepares for data analysis.
[0311] Step 3:
[0312] The server uses a generative AI model to convert speech data into text. It uses Google Speech-to-Text to generate corresponding strings from the speech data and extract urgent keywords. The input is speech data, and the output is textual information. In this process, the user's spoken content is converted into specific textual information.
[0313] Step 4:
[0314] The server uses image analysis techniques such as OpenCV to analyze video data and evaluate the environmental conditions and changes. The input is video data, and the output is feature information of the analyzed video. In this step, high-risk elements within the video are identified.
[0315] Step 5:
[0316] The server integrates the results of voice-to-text and video analysis to determine the priority of notifications based on urgency. This process prioritizes notifications by considering emergency keywords and environmental analysis information. The input is transcribed voice information and video feature information, and the output is prioritized notification data.
[0317] Step 6:
[0318] The server notifies the appropriate emergency services of prioritized alert data. The input is prioritized alert information, which becomes the content of the notification to the emergency services. The server refers to the database, selects the most suitable service based on the content of each piece of information, and sends a notification to encourage a rapid response.
[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0320] This invention relates to an embodiment of an emergency notification system for disasters and emergencies that incorporates an emotion engine to recognize the user's emotions. In this embodiment, the user inputs an emergency message using a dedicated application on their smartphone. The terminal collects emotion data along with voice and image data and transmits it to a server via satellite communication.
[0321] The transmitted data is processed on the server using an analysis method that utilizes generation AI. Within this process, an emotion engine recognizes the user's emotions through voice analysis, and this information is used to determine the level of urgency. For example, based on changes in the user's voice tone and volume, features indicating heightened emotions are extracted and considered as factors that increase the level of urgency.
[0322] Furthermore, image analysis is performed to determine the environment and the extent of injuries from visual information. This information is integrated on a server, and a prioritization mechanism determines the priority of the response. The emotion engine evaluates the impact of emotional information on the severity and urgency of the report by comparing it with past databases.
[0323] Next, the server notifies emergency services of the prioritized information. This process also includes emotional data, enabling emergency services to respond appropriately, taking into account the psychological state of the victims. For example, based on emotional information, emergency services can instruct immediate action on cases that require priority.
[0324] Ultimately, a feedback mechanism improves the accuracy of the operational system through feedback from emergency services. Based on this feedback, the emotion engine algorithm is improved, and the system is continuously adjusted to achieve even greater accuracy in future incidents. This results in a comprehensive emergency call system that enables rapid and highly accurate responses.
[0325] The following describes the processing flow.
[0326] Step 1:
[0327] The user enters an emergency call message and records a voice message through a smartphone application. The device inputs the user's voice into an emotion engine in real time and acquires data to estimate their emotions.
[0328] Step 2:
[0329] The device packages the acquired voice messages, emotion data, and location information and transmits them to the server via satellite communication. The transmitted information also includes the results of the user's emotion analysis.
[0330] Step 3:
[0331] The server inputs the received data into an analysis module and uses a generation AI to convert the voice message into text. Furthermore, it evaluates the user's psychological urgency based on the emotional data provided by the emotion engine.
[0332] Step 4:
[0333] The server analyzes the collected image data using image analysis algorithms to extract information about the damage and the environment. It then references recognized emotion data to comprehensively assess the overall urgency of the situation.
[0334] Step 5:
[0335] The server centralizes the analyzed urgency and sentiment data and uses prioritization mechanisms to determine the priority of rescue operations. Based on this information, cases requiring a rapid response are identified.
[0336] Step 6:
[0337] The server notifies emergency services of priority information. This notification includes detailed disaster information as well as information about the user's emotional state, allowing emergency services to accurately understand the situation and take appropriate action when psychological support is needed.
[0338] Step 7:
[0339] The server continuously monitors feedback from emergency calls and adjusts its emotion engine and analysis methods based on reports from emergency services. This improves the system so that it can provide a more accurate emergency response in the future.
[0340] (Example 2)
[0341] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0342] In the event of a disaster or emergency, accurate analysis of on-site reports is essential for a swift and appropriate response. However, conventional systems lack the accuracy to analyze voice and image data, making it particularly difficult to determine the level of urgency while considering the user's emotional state. Furthermore, coordinating with emergency services presents challenges in responding to the specific needs of the field.
[0343] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0344] In this invention, the server includes a communication device that receives emergency calls using satellite communication, an analysis mechanism that converts voice signals into text data and analyzes image information using generation AI technology, and an emergency assessment device that recognizes the user's emotions and evaluates the urgency by integrating that emotion information with the analyzed data. This makes it possible to accurately determine the priority of responses in the event of a disaster or distress and to provide a swift and appropriate emergency response.
[0345] A "communication device" is a device that uses satellite communication to receive emergency calls from outside.
[0346] An "analysis mechanism" is a device that uses generative AI technology to convert audio signals into text data or to perform detailed analysis of image information.
[0347] An "urgency assessment device" is a device that analyzes the user's emotions, evaluates the local situation along with the integrated data, and determines the level of urgency.
[0348] A "prioritization device" is a device that determines the priority of responses based on the urgency level obtained by an urgency assessment device.
[0349] A "communication device" is a device used to accurately and quickly notify relevant organizations of emergency information.
[0350] A "regulating device" is a device used to improve the analysis mechanism based on feedback from related organizations, thereby enhancing the overall accuracy of the system.
[0351] This invention is a system for effectively making emergency calls during disasters or emergencies. The system consists of a mechanism in which a user inputs an emergency message using a smartphone terminal, and this information is transmitted to a server via satellite communication.
[0352] The device is equipped with the functionality to collect voice and image data. Users use a dedicated smartphone application to make voice notifications and send images. The device utilizes speech recognition technology to convert the user's voice into text data. This conversion uses a generative AI model, providing more advanced speech analysis capabilities.
[0353] The server receives data transmitted from the terminal and processes it using an analysis mechanism. This analysis mechanism, which includes a generative AI, not only converts audio data into text but also has the ability to identify emotions from factors such as tone and volume. Furthermore, it can evaluate the user's visual situation using image data.
[0354] For example, consider a scenario where a user is lost in a mountainous area and uses their device to record a voice message saying, "Please help me, I don't know where I am," and sends it along with an image. An example of a prompt message in this case would be, "We have received your emergency call, will analyze the user's emotions from the voice data, and will develop a rapid response plan."
[0355] The server's urgency assessment device quickly determines the urgency of the situation based on analyzed emotional and visual information. Based on this, the priority determination device organizes the response priorities and notifies those with the highest level of urgency first.
[0356] Finally, the server notifies the relevant organizations of appropriate response information through a communication device. The communication device conveys detailed information about the report, including the user's emotional state, to the relevant organizations, helping them to implement the most appropriate response on the scene.
[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0358] Step 1:
[0359] The user enters an emergency message using a dedicated smartphone application and records audio and image data to the device. This input includes the user's voice and video captured by the smartphone camera. The device formats this data and saves it in an appropriate format for subsequent processing.
[0360] Step 2:
[0361] The device processes recorded audio data and converts it into text data using a generation AI model. The input is the user's audio file, which is then passed through an analysis engine to convert it into text information. This outputs specific textual information. The intonation and tone of the voice are also analyzed simultaneously, and characteristic data indicating emotion is extracted.
[0362] Step 3:
[0363] The device analyzes stored image data and extracts visual features. The input is an image file taken by the user, and an image processing algorithm is used to evaluate the environmental conditions and the state of people. The output includes visual hazards and environmental information derived from the image.
[0364] Step 4:
[0365] The terminal aggregates the converted text data, sentiment data, and image analysis results and sends them to the server via the internet. This transmission process involves data integration and packetization to ensure secure and rapid transmission to the server.
[0366] Step 5:
[0367] The server processes the received data using an analysis mechanism. It analyzes text data to identify urgent elements and evaluates the user's psychological state based on emotional data extracted from the audio. This process efficiently identifies information with high urgency.
[0368] Step 6:
[0369] The server integrates the emotional and image information analyzed by the urgency assessment device to calculate an urgency score. In this step, the urgency of each input data is quantified and prioritized. The final output is a response plan with determined priorities.
[0370] Step 7:
[0371] The server notifies relevant organizations of prioritized information via communication devices. This notification includes the user's location and instructions for responding to high-priority cases. This allows for immediate and appropriate responses on the ground.
[0372] (Application Example 2)
[0373] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0374] Conventional emergency reporting systems often judge the urgency of a situation based on a uniform standard without considering the caller's psychological state or emotions, which can lead to a lack of prompt and appropriate responses tailored to the actual situation. As a result, there is a problem where the prioritization of emergency responses is incorrect, and immediate responses may be delayed. Furthermore, responding without considering the emotional state of the victims can increase their psychological burden.
[0375] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0376] In this invention, the server includes means for receiving emergency calls using satellite communication, means for converting voice information into text using a generation AI and analyzing video information, means for aggregating the analyzed information and determining a ranking based on the degree of urgency, and means for analyzing the user's emotions using an emotion engine and considering this in determining the degree of urgency. This makes it possible to evaluate the overall degree of urgency, including the emotional state of the caller, enabling faster and more appropriate prioritization and response.
[0377] "Satellite communication" is a means of communication that uses artificial satellites to send and receive information between devices located in distant places on Earth.
[0378] An "emergency call receiving device" is a device that quickly receives emergency messages from callers in the event of an abnormal situation or danger.
[0379] "Generative AI" refers to artificial intelligence technology that can analyze various types of data, including audio and video, and extract useful information from them.
[0380] "Auditory information" refers to data that records the words and sounds spoken by a person, and it serves as the basis for analyzing context and emotions.
[0381] "Visual information" refers to visual data acquired through cameras and other imaging devices, and is information that visually captures the state of objects and the environment.
[0382] "Analysis methods" refer to technical techniques for processing collected data and extracting meaningful information from it.
[0383] A "system for determining priority based on urgency" is a system for evaluating the seriousness and urgency of a situation and determining the priority of responses accordingly.
[0384] An "emotion engine" is a technological element that can identify and analyze a person's emotional state from audio and video.
[0385] "The caller's emotional state" refers to the feelings and psychological condition of the person making the emergency call.
[0386] "User" refers to the individual or organization that uses this emergency call system.
[0387] The system that realizes this invention consists of the cooperation of multiple devices and software. When making an emergency call, the user uses a portable information terminal. This terminal is equipped with satellite communication capabilities, enabling highly reliable communication unaffected by the external environment.
[0388] When a user presses the emergency button, the device's microphone records audio and the camera captures video. This information is processed in real time by a generative AI; the audio is converted to text and the video is analyzed. The generative AI uses the Google Cloud Speech-to-Text API to transcribe the audio, while the Google Cloud Vision API is used to analyze the video. This enables detailed analysis of both audio and video.
[0389] The server receives information transmitted from the terminal and analyzes the user's emotions using an emotion engine. In this process, it detects changes in the tone and volume of the user's voice to determine the level of emotional arousal. It also compares this data with past data to determine the degree of urgency. The server sets priorities according to the degree of urgency and immediately notifies emergency services if necessary.
[0390] For example, if a user witnesses someone behaving suspiciously on the street, they can report the emergency situation through the application. Information reflecting the urgency of the voice report is immediately sent to the server, and if action is needed, instructions are quickly given to emergency services. Through this process, it is possible to ensure an appropriate and rapid response in emergencies.
[0391] An example of a prompt to smoothly initiate the analysis of a generative AI model would be: "Based on recent emergency call cases, analyze the emotions in this voice data. Provide guidance to assess the urgency and select the appropriate response route." This allows for more accurate assessment of emotional changes and urgency.
[0392] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0393] Step 1:
[0394] When a user presses the emergency call button on their mobile device, the device begins recording audio and capturing video with its camera. The input consists of the user's audio and video data, which is then converted into a format that can be processed by the AI model. Specifically, this involves collecting the necessary data using the device's microphone and camera.
[0395] Step 2:
[0396] The device uses generative AI to convert recorded audio information into text. This involves analyzing the audio data using the Google Cloud Speech-to-Text API and converting it into text. The input is audio data, and the output is text data. The process includes extracting audio features within the device and processing them as text strings.
[0397] Step 3:
[0398] The device analyzes captured video information using the Google Cloud Vision API. The input is video data, and the output is information derived from the analysis of various visual elements. Specifically, it recognizes objects and environmental characteristics within the image and collects this information as data.
[0399] Step 4:
[0400] The terminal transmits the acquired text information and video analysis results to the server. This is done using satellite communication, ensuring highly reliable data transfer. The input is the transcribed audio information and analysis results, and the output is the information sent to the server.
[0401] Step 5:
[0402] The server analyzes the user's emotions using an emotion engine based on the information it receives. For example, it can infer stress and tension from factors such as voice tone and volume. The input is the user's voice tone information, and the output is the analyzed emotion information. The generation of emotion data within the server is a specific operation.
[0403] Step 6:
[0404] The server aggregates the analysis results and sets priorities based on urgency. During this process, the analysis results are compared with other emergency call data to determine which response should be given the highest priority. The input consists of analyzed emotional information and video data, and the output is the priority ranking. The main operation involves data comparison processing performed internally by the server.
[0405] Step 7:
[0406] Based on the priority determined by the server, a notification is sent to emergency services. This notification includes audio data, video data, and an analyzed level of urgency, enabling an appropriate response. The input is the data of the set priority, and the output is the notification to emergency services. The automatic transmission of data from the server is the specific operation in this processing step.
[0407] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0408] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0409] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0410] [Third Embodiment]
[0411] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0412] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0413] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0414] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0415] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0417] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0418] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0419] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0420] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0421] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0422] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0423] This invention relates to an embodiment of an emergency notification system utilizing satellite communication. When a user sends an SOS message using their smartphone during a disaster or emergency, the system receives the data via satellite communication. The received data is then processed by a server.
[0424] The server first uses a generation AI to convert audio data into text. Machine learning techniques are used to extract keywords indicating urgency from the audio and record them as text. Simultaneously, image analysis algorithms are used to analyze the situation of disaster victims and changes in the environment from photos and videos. For example, if a user's image shows collapsed buildings or injured people, the server identifies them and calculates the urgency as further information.
[0425] The server then aggregates the analyzed data and determines a priority for each emergency call. It selects the most important calls from a large number and sorts them in order of the need for immediate response. This process allows for an efficient response to situations requiring emergency attention. For example, calls from seriously injured individuals or high-risk areas are prioritized within the system, and rescue operations are coordinated to ensure rapid response.
[0426] Prioritized information is sent from the server to emergency services. The server then refers to a database of pre-registered emergency services and sends the necessary information to the most appropriate service provider. This notification includes the user's location, the extent of the damage, and recommended emergency responses, enabling the emergency services to make informed decisions.
[0427] As the final step in this embodiment, the server optimizes the system based on feedback received after the notification. Based on the feedback provided by the emergency services, the system can make real-time adjustments to further improve accuracy. This improves the speed and accuracy of responses in similar cases.
[0428] Because this system does not rely on conventional terrestrial communications, it minimizes the impact of communication disruptions during disasters and enables the rapid rescue of victims.
[0429] The following describes the processing flow.
[0430] Step 1:
[0431] When a user creates an emergency call or SOS message using a dedicated application on their smartphone and presses the send button, the message is sent. At this point, the user's location information and message content are transmitted from the device via satellite communication.
[0432] Step 2:
[0433] The server acquires emergency call data received via satellite communication. Immediately after reception, the data is transferred to an analysis module. Here, the audio and image data are ready to be processed separately.
[0434] Step 3:
[0435] The server uses a generation AI to convert voice data into text. Using speech recognition technology, it extracts keywords and phrases indicating urgency, which serve as the basis for determining the urgency of the report. Simultaneously, it uses an analysis algorithm to analyze the environment and injury status from image data, obtaining information necessary for accurate situation assessment.
[0436] Step 4:
[0437] The server centralizes the analyzed information and prioritizes each report. An AI algorithm automatically prioritizes rescue cases based on the urgency and location of the victims. Furthermore, detailed disaster patterns based on the user's equipment and surrounding environment are simultaneously analyzed.
[0438] Step 5:
[0439] The server promptly notifies emergency services. The notification includes special details about the disaster, location information, and urgency ranking, providing data that enables immediate response. The notification protocol is automatically configured based on a pre-registered database of emergency services.
[0440] Step 6:
[0441] The server continuously monitors and gathers feedback after an alert is received. Based on feedback from emergency services, the system is continuously adjusted, and AI algorithms are improved and databases are updated in real time. Through this process, the overall accuracy and efficiency of the system's response are enhanced.
[0442] (Example 1)
[0443] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0444] During disasters and emergencies, there is a need to ensure that victims can receive emergency assistance quickly and effectively, even when ground-based communication infrastructure is unavailable. Conventional systems have the problem of hindering the transmission of emergency calls due to communication disruptions. Furthermore, there is a challenge in that there are insufficient means to accurately assess the urgency of received information and quickly coordinate with the appropriate emergency services, leading to delays in response.
[0445] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0446] In this invention, the server includes a device that receives emergency calls using satellite communication, a device that converts voice information into text and analyzes image information using generative artificial intelligence, and a device that aggregates the analyzed information and determines the processing order based on the urgency. This makes it possible to quickly determine the urgency and transmit information to the appropriate emergency agency even if ground communication is interrupted.
[0447] "Satellite communication" is a technology that transmits and receives information via communication satellites without using terrestrial communication infrastructure.
[0448] An "emergency call" is communication information sent to request help in emergency situations such as disasters or getting lost.
[0449] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to analyze data and generate information such as speech and text.
[0450] "Converting audio information to text" is the process of analyzing audio and representing it as a corresponding string of characters.
[0451] "Analyzing image information" is a technique that uses image data to identify situations and objects and extract meaningful information.
[0452] "Device" refers to hardware or software designed to perform a series of processes.
[0453] "Urgency" is an indicator that shows the degree to which an emergency response is necessary, and it is a criterion for determining the order in which a rapid response is required.
[0454] "Processing order" refers to guidelines for determining the order in which multiple tasks are performed based on their priority.
[0455] An "emergency agency" is a public or private organization that has the facilities and personnel to respond to emergencies such as disasters and accidents.
[0456] This invention is an emergency notification system that utilizes satellite communication to enable users to receive rapid assistance in emergency situations.
[0457] First, the user uses a dedicated smartphone app. This app is used to send SOS messages in the event of a disaster or emergency. The device is equipped with a satellite communication module and transmits signals via an airborne communication network without using ground-based communication infrastructure. A general-purpose module that supports various protocols and standards could be used as the satellite communication module.
[0458] The server processes data received via satellite communication. Generative artificial intelligence is used to transcribe audio data into text. A general-purpose AI model that excels at speech recognition and can perform effective text conversion is considered. Next, machine learning algorithms are applied to the text data to extract specific keywords indicating urgency. For image data, commonly used image analysis algorithms are used to determine the disaster situation from the video.
[0459] The analyzed data is aggregated as an urgency score, and the processing order is determined based on its importance. This information is notified to the nearest emergency services, and instructions are given to begin immediate action. This process enables rapid and accurate emergency assistance.
[0460] For example, if a user makes a voice call saying, "My house is on fire, please help," the voice data is converted to text by the server, and keywords such as "fire" and "help" are extracted as factors that increase the urgency. As a result, the server processes it as a high-priority call and can immediately transmit the information to fire departments and other relevant organizations.
[0461] Example of a prompt:
[0462] "A user has sent an SOS message from a disaster area. Analyze the audio and video, calculate the urgency, and notify the appropriate agency."
[0463] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0464] Step 1:
[0465] Procedure for users to send SOS messages
[0466] Users launch a dedicated app on their smartphones to record newly occurring emergencies. A voice input function allows them to enter comments describing the situation using voice. They can also capture images and videos of the scene using the camera function and attach them to the SOS message. This data is transmitted from the device to the server via a satellite communication module. Input consists of voice and image data, while output is structured data sent to the server.
[0467] Step 2:
[0468] The server receives the data and converts the audio data into text.
[0469] The server receives the SOS message sent from the terminal. The received audio data is converted into text using a generative AI model. Specifically, the audio waveform data is input to the AI model, which is then analyzed and output as text data.
[0470] Step 3:
[0471] The server extracts urgent keywords from the text.
[0472] The server analyzes the text data and extracts keywords indicating urgency (e.g., "Help," "Fire"). It applies a machine learning algorithm to identify important keywords from a pre-trained model. The input is text data, and the output is the extracted urgency keywords.
[0473] Step 4:
[0474] The process by which the server analyzes image data.
[0475] The server analyzes the received image data to determine the situation and environment of the subject. Using image analysis algorithms, it identifies, for example, collapsed buildings or the condition of people. The input is image data, and the output is the analyzed situational information.
[0476] Step 5:
[0477] A procedure in which the server calculates urgency and determines the processing order.
[0478] The server integrates keywords from voice analysis with image analysis results to calculate the overall urgency. It generates an urgency score using a proprietary calculation model and compares it to other reports based on that score. The output is this score, which is used to determine the processing order.
[0479] Step 6:
[0480] The server notifies the most appropriate emergency agency.
[0481] Based on the determined processing order, the server notifies the most appropriate emergency agency. The information sent to each emergency agency includes the user's location information, transcribed audio information, and analyzed image information. This allows the receiving agency to quickly understand the situation and begin responding. The output is notification data.
[0482] (Application Example 1)
[0483] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0484] In natural disasters and emergencies, there is a need for communication methods that enable rapid and accurate reporting and initial response. However, conventional systems that rely on terrestrial communications struggle to effectively report emergencies when communication infrastructure is destroyed. Furthermore, they lack intuitive operation methods to enable users to report emergencies at the appropriate time. In addition, there are few means to effectively determine the priority of emergency calls, hindering efficient rescue operations.
[0485] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0486] In this invention, the server includes a receiving means for receiving emergency calls using satellite communication, an analysis means for converting voice data into text using a generation AI and analyzing image data, and a prioritizing means for aggregating the analyzed data and determining priorities based on urgency. This realizes an emergency call system that does not depend on communication infrastructure, allows users to make calls intuitively, and enables rapid rescue operations through efficient prioritization.
[0487] "Satellite communication" is a technology that transmits and receives communication signals via artificial satellites, making it possible to transmit signals over a wide area without relying on ground-based communication infrastructure.
[0488] "Generative AI" is an artificial intelligence technology that automatically generates information from data, and has the ability to analyze various types of data such as audio and images to extract text and features.
[0489] "Converting audio data to text" is the process of analyzing audio information and converting it into corresponding textual information.
[0490] "Analyzing image data" refers to the technique of extracting information from images and identifying specific features or patterns.
[0491] "Determining priorities" means evaluating the importance of various data and reports and setting the order in which they require immediate attention.
[0492] "Recording means installed in a portable display device" refers to a function incorporated into a display device that can be easily carried by an individual, which acquires and records audio and video.
[0493] A "startup mechanism" is a function that automatically starts the system when a user performs an action that meets specific conditions.
[0494] A "notification mechanism" is a function for appropriately transmitting analyzed information to relevant authorities and agencies.
[0495] To realize this invention, a portable display device carried by the user, a server, and a satellite communication network are required. The portable display device includes a microphone and camera necessary for acquiring audio and video. When the user gives a voice command in an emergency, the device automatically starts recording audio and video data. This data is transmitted from the device to the server via satellite communication.
[0496] The server uses a generative AI model to convert audio data into text and extracts keywords from that text to assess urgency. Specifically, it uses Google Speech-to-Text to convert audio to text and then analyzes the extracted text with TensorFlow. Additionally, it uses image processing technologies such as OpenCV to evaluate environmental changes and risk levels in video data.
[0497] Based on the analyzed data, the server prioritizes emergency events and generates a list of notifications, ordered from those requiring immediate attention to those requiring it. This enables efficient rescue operations. By utilizing satellite communication, rather than relying on existing communication infrastructure, a highly reliable notification system can be established even in the event of communication disruptions.
[0498] As a concrete example, in a mountainous area where ground communication is cut off due to a natural disaster, if a hiker is about to fall, they can simply say "SOS" into a portable display device, and the appropriate agency will be quickly notified via satellite communication.
[0499] An example of a prompt message for a generative AI model to analyze audio data is: "Design a machine learning model to analyze audio data and extract urgency keywords. Also, propose a method for evaluating the level of risk through image analysis."
[0500] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0501] Step 1:
[0502] The user speaks "SOS" into the portable display device to report an emergency. The device records this audio as sound data and simultaneously acquires video data using its camera. The audio and video inputs are used to make the emergency call.
[0503] Step 2:
[0504] The terminal transmits acquired audio and video data to the server via satellite communication. The input is data recorded within the terminal, which becomes the information to be transmitted to the server. The server receives this data and prepares for data analysis.
[0505] Step 3:
[0506] The server uses a generative AI model to convert speech data into text. It uses Google Speech-to-Text to generate corresponding strings from the speech data and extract urgent keywords. The input is speech data, and the output is textual information. In this process, the user's spoken content is converted into specific textual information.
[0507] Step 4:
[0508] The server uses image analysis techniques such as OpenCV to analyze video data and evaluate the environmental conditions and changes. The input is video data, and the output is feature information of the analyzed video. In this step, high-risk elements within the video are identified.
[0509] Step 5:
[0510] The server integrates the results of voice-to-text and video analysis to determine the priority of notifications based on urgency. This process prioritizes notifications by considering emergency keywords and environmental analysis information. The input is transcribed voice information and video feature information, and the output is prioritized notification data.
[0511] Step 6:
[0512] The server notifies the appropriate emergency services of prioritized alert data. The input is prioritized alert information, which becomes the content of the notification to the emergency services. The server refers to the database, selects the most suitable service based on the content of each piece of information, and sends a notification to encourage a rapid response.
[0513] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0514] This invention relates to an embodiment of an emergency notification system for disasters and emergencies that incorporates an emotion engine to recognize the user's emotions. In this embodiment, the user inputs an emergency message using a dedicated application on their smartphone. The terminal collects emotion data along with voice and image data and transmits it to a server via satellite communication.
[0515] The transmitted data is processed on the server using an analysis method that utilizes generation AI. Within this process, an emotion engine recognizes the user's emotions through voice analysis, and this information is used to determine the level of urgency. For example, based on changes in the user's voice tone and volume, features indicating heightened emotions are extracted and considered as factors that increase the level of urgency.
[0516] Furthermore, image analysis is performed to determine the environment and the extent of injuries from visual information. This information is integrated on a server, and a prioritization mechanism determines the priority of the response. The emotion engine evaluates the impact of emotional information on the severity and urgency of the report by comparing it with past databases.
[0517] Next, the server notifies emergency services of the prioritized information. This process also includes emotional data, enabling emergency services to respond appropriately, taking into account the psychological state of the victims. For example, based on emotional information, emergency services can instruct immediate action on cases that require priority.
[0518] Ultimately, a feedback mechanism improves the accuracy of the operational system through feedback from emergency services. Based on this feedback, the emotion engine algorithm is improved, and the system is continuously adjusted to achieve even greater accuracy in future incidents. This results in a comprehensive emergency call system that enables rapid and highly accurate responses.
[0519] The following describes the processing flow.
[0520] Step 1:
[0521] The user enters an emergency call message and records a voice message through a smartphone application. The device inputs the user's voice into an emotion engine in real time and acquires data to estimate their emotions.
[0522] Step 2:
[0523] The device packages the acquired voice messages, emotion data, and location information and transmits them to the server via satellite communication. The transmitted information also includes the results of the user's emotion analysis.
[0524] Step 3:
[0525] The server inputs the received data into an analysis module and uses a generation AI to convert the voice message into text. Furthermore, it evaluates the user's psychological urgency based on the emotional data provided by the emotion engine.
[0526] Step 4:
[0527] The server analyzes the collected image data using image analysis algorithms to extract information about the damage and the environment. It then references recognized emotion data to comprehensively assess the overall urgency of the situation.
[0528] Step 5:
[0529] The server centralizes the analyzed urgency and sentiment data and uses prioritization mechanisms to determine the priority of rescue operations. Based on this information, cases requiring a rapid response are identified.
[0530] Step 6:
[0531] The server notifies emergency services of priority information. This notification includes detailed disaster information as well as information about the user's emotional state, allowing emergency services to accurately understand the situation and take appropriate action when psychological support is needed.
[0532] Step 7:
[0533] The server continuously monitors feedback from emergency calls and adjusts its emotion engine and analysis methods based on reports from emergency services. This improves the system so that it can provide a more accurate emergency response in the future.
[0534] (Example 2)
[0535] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] In the event of a disaster or emergency, accurate analysis of on-site reports is essential for a swift and appropriate response. However, conventional systems lack the accuracy to analyze voice and image data, making it particularly difficult to determine the level of urgency while considering the user's emotional state. Furthermore, coordinating with emergency services presents challenges in responding to the specific needs of the field.
[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0538] In this invention, the server includes a communication device that receives emergency calls using satellite communication, an analysis mechanism that converts voice signals into text data and analyzes image information using generation AI technology, and an emergency assessment device that recognizes the user's emotions and evaluates the urgency by integrating that emotion information with the analyzed data. This makes it possible to accurately determine the priority of responses in the event of a disaster or distress and to provide a swift and appropriate emergency response.
[0539] A "communication device" is a device that uses satellite communication to receive emergency calls from outside.
[0540] An "analysis mechanism" is a device that uses generative AI technology to convert audio signals into text data or to perform detailed analysis of image information.
[0541] An "urgency assessment device" is a device that analyzes the user's emotions, evaluates the local situation along with the integrated data, and determines the level of urgency.
[0542] A "prioritization device" is a device that determines the priority of responses based on the urgency level obtained by an urgency assessment device.
[0543] A "communication device" is a device used to accurately and quickly notify relevant organizations of emergency information.
[0544] A "regulating device" is a device used to improve the analysis mechanism based on feedback from related organizations, thereby enhancing the overall accuracy of the system.
[0545] This invention is a system for effectively making emergency calls during disasters or emergencies. The system consists of a mechanism in which a user inputs an emergency message using a smartphone terminal, and this information is transmitted to a server via satellite communication.
[0546] The device is equipped with the functionality to collect voice and image data. Users use a dedicated smartphone application to make voice notifications and send images. The device utilizes speech recognition technology to convert the user's voice into text data. This conversion uses a generative AI model, providing more advanced speech analysis capabilities.
[0547] The server receives data transmitted from the terminal and processes it using an analysis mechanism. This analysis mechanism, which includes a generative AI, not only converts audio data into text but also has the ability to identify emotions from factors such as tone and volume. Furthermore, it can evaluate the user's visual situation using image data.
[0548] For example, consider a scenario where a user is lost in a mountainous area and uses their device to record a voice message saying, "Please help me, I don't know where I am," and sends it along with an image. An example of a prompt message in this case would be, "We have received your emergency call, will analyze the user's emotions from the voice data, and will develop a rapid response plan."
[0549] The server's urgency assessment device quickly determines the urgency of the situation based on analyzed emotional and visual information. Based on this, the priority determination device organizes the response priorities and notifies those with the highest level of urgency first.
[0550] Finally, the server notifies the relevant organizations of appropriate response information through a communication device. The communication device conveys detailed information about the report, including the user's emotional state, to the relevant organizations, helping them to implement the most appropriate response on the scene.
[0551] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0552] Step 1:
[0553] The user enters an emergency message using a dedicated smartphone application and records audio and image data to the device. This input includes the user's voice and video captured by the smartphone camera. The device formats this data and saves it in an appropriate format for subsequent processing.
[0554] Step 2:
[0555] The device processes recorded audio data and converts it into text data using a generation AI model. The input is the user's audio file, which is then passed through an analysis engine to convert it into text information. This outputs specific textual information. The intonation and tone of the voice are also analyzed simultaneously, and characteristic data indicating emotion is extracted.
[0556] Step 3:
[0557] The device analyzes stored image data and extracts visual features. The input is an image file taken by the user, and an image processing algorithm is used to evaluate the environmental conditions and the state of people. The output includes visual hazards and environmental information derived from the image.
[0558] Step 4:
[0559] The terminal aggregates the converted text data, sentiment data, and image analysis results and sends them to the server via the internet. This transmission process involves data integration and packetization to ensure secure and rapid transmission to the server.
[0560] Step 5:
[0561] The server processes the received data using an analysis mechanism. It analyzes text data to identify urgent elements and evaluates the user's psychological state based on emotional data extracted from the audio. This process efficiently identifies information with high urgency.
[0562] Step 6:
[0563] The server integrates the emotional and image information analyzed by the urgency assessment device to calculate an urgency score. In this step, the urgency of each input data is quantified and prioritized. The final output is a response plan with determined priorities.
[0564] Step 7:
[0565] The server notifies relevant organizations of prioritized information via communication devices. This notification includes the user's location and instructions for responding to high-priority cases. This allows for immediate and appropriate responses on the ground.
[0566] (Application Example 2)
[0567] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0568] Conventional emergency reporting systems often judge the urgency of a situation based on a uniform standard without considering the caller's psychological state or emotions, which can lead to a lack of prompt and appropriate responses tailored to the actual situation. As a result, there is a problem where the prioritization of emergency responses is incorrect, and immediate responses may be delayed. Furthermore, responding without considering the emotional state of the victims can increase their psychological burden.
[0569] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0570] In this invention, the server includes means for receiving emergency calls using satellite communication, means for converting voice information into text using a generation AI and analyzing video information, means for aggregating the analyzed information and determining a ranking based on the degree of urgency, and means for analyzing the user's emotions using an emotion engine and considering this in determining the degree of urgency. This makes it possible to evaluate the overall degree of urgency, including the emotional state of the caller, enabling faster and more appropriate prioritization and response.
[0571] "Satellite communication" is a means of communication that uses artificial satellites to send and receive information between devices located in distant places on Earth.
[0572] An "emergency call receiving device" is a device that quickly receives emergency messages from callers in the event of an abnormal situation or danger.
[0573] "Generative AI" refers to artificial intelligence technology that can analyze various types of data, including audio and video, and extract useful information from them.
[0574] "Auditory information" refers to data that records the words and sounds spoken by a person, and it serves as the basis for analyzing context and emotions.
[0575] "Visual information" refers to visual data acquired through cameras and other imaging devices, and is information that visually captures the state of objects and the environment.
[0576] "Analysis methods" refer to technical techniques for processing collected data and extracting meaningful information from it.
[0577] A "system for determining priority based on urgency" is a system for evaluating the seriousness and urgency of a situation and determining the priority of responses accordingly.
[0578] An "emotion engine" is a technological element that can identify and analyze a person's emotional state from audio and video.
[0579] "The caller's emotional state" refers to the feelings and psychological condition of the person making the emergency call.
[0580] "User" refers to the individual or organization that uses this emergency call system.
[0581] The system that realizes this invention consists of the cooperation of multiple devices and software. When making an emergency call, the user uses a portable information terminal. This terminal is equipped with satellite communication capabilities, enabling highly reliable communication unaffected by the external environment.
[0582] When a user presses the emergency button, the device's microphone records audio and the camera captures video. This information is processed in real time by a generative AI; the audio is converted to text and the video is analyzed. The generative AI uses the Google Cloud Speech-to-Text API to transcribe the audio, while the Google Cloud Vision API is used to analyze the video. This enables detailed analysis of both audio and video.
[0583] The server receives information transmitted from the terminal and analyzes the user's emotions using an emotion engine. In this process, it detects changes in the tone and volume of the user's voice to determine the level of emotional arousal. It also compares this data with past data to determine the degree of urgency. The server sets priorities according to the degree of urgency and immediately notifies emergency services if necessary.
[0584] For example, if a user witnesses someone behaving suspiciously on the street, they can report the emergency situation through the application. Information reflecting the urgency of the voice report is immediately sent to the server, and if action is needed, instructions are quickly given to emergency services. Through this process, it is possible to ensure an appropriate and rapid response in emergencies.
[0585] An example of a prompt to smoothly initiate the analysis of a generative AI model would be: "Based on recent emergency call cases, analyze the emotions in this voice data. Provide guidance to assess the urgency and select the appropriate response route." This allows for more accurate assessment of emotional changes and urgency.
[0586] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0587] Step 1:
[0588] When a user presses the emergency call button on their mobile device, the device begins recording audio and capturing video with its camera. The input consists of the user's audio and video data, which is then converted into a format that can be processed by the AI model. Specifically, this involves collecting the necessary data using the device's microphone and camera.
[0589] Step 2:
[0590] The device uses generative AI to convert recorded audio information into text. This involves analyzing the audio data using the Google Cloud Speech-to-Text API and converting it into text. The input is audio data, and the output is text data. The process includes extracting audio features within the device and processing them as text strings.
[0591] Step 3:
[0592] The device analyzes captured video information using the Google Cloud Vision API. The input is video data, and the output is information derived from the analysis of various visual elements. Specifically, it recognizes objects and environmental characteristics within the image and collects this information as data.
[0593] Step 4:
[0594] The terminal transmits the acquired text information and video analysis results to the server. This is done using satellite communication, ensuring highly reliable data transfer. The input is the transcribed audio information and analysis results, and the output is the information sent to the server.
[0595] Step 5:
[0596] The server analyzes the user's emotions using an emotion engine based on the information it receives. For example, it can infer stress and tension from factors such as voice tone and volume. The input is the user's voice tone information, and the output is the analyzed emotion information. The generation of emotion data within the server is a specific operation.
[0597] Step 6:
[0598] The server aggregates the analysis results and sets priorities based on urgency. During this process, the analysis results are compared with other emergency call data to determine which response should be given the highest priority. The input consists of analyzed emotional information and video data, and the output is the priority ranking. The main operation involves data comparison processing performed internally by the server.
[0599] Step 7:
[0600] Based on the priority determined by the server, a notification is sent to emergency services. This notification includes audio data, video data, and an analyzed level of urgency, enabling an appropriate response. The input is the data of the set priority, and the output is the notification to emergency services. The automatic transmission of data from the server is the specific operation in this processing step.
[0601] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0602] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0603] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0604] [Fourth Embodiment]
[0605] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0606] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0607] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0608] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0609] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0610] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0611] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0612] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0613] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0614] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0615] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0616] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0617] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0618] This invention relates to an embodiment of an emergency notification system utilizing satellite communication. When a user sends an SOS message using their smartphone during a disaster or emergency, the system receives the data via satellite communication. The received data is then processed by a server.
[0619] The server first uses a generation AI to convert audio data into text. Machine learning techniques are used to extract keywords indicating urgency from the audio and record them as text. Simultaneously, image analysis algorithms are used to analyze the situation of disaster victims and changes in the environment from photos and videos. For example, if a user's image shows collapsed buildings or injured people, the server identifies them and calculates the urgency as further information.
[0620] The server then aggregates the analyzed data and determines a priority for each emergency call. It selects the most important calls from a large number and sorts them in order of the need for immediate response. This process allows for an efficient response to situations requiring emergency attention. For example, calls from seriously injured individuals or high-risk areas are prioritized within the system, and rescue operations are coordinated to ensure rapid response.
[0621] Prioritized information is sent from the server to emergency services. The server then refers to a database of pre-registered emergency services and sends the necessary information to the most appropriate service provider. This notification includes the user's location, the extent of the damage, and recommended emergency responses, enabling the emergency services to make informed decisions.
[0622] As the final step in this embodiment, the server optimizes the system based on feedback received after the notification. Based on the feedback provided by the emergency services, the system can make real-time adjustments to further improve accuracy. This improves the speed and accuracy of responses in similar cases.
[0623] Because this system does not rely on conventional terrestrial communications, it minimizes the impact of communication disruptions during disasters and enables the rapid rescue of victims.
[0624] The following describes the processing flow.
[0625] Step 1:
[0626] When a user creates an emergency call or SOS message using a dedicated application on their smartphone and presses the send button, the message is sent. At this point, the user's location information and message content are transmitted from the device via satellite communication.
[0627] Step 2:
[0628] The server acquires emergency call data received via satellite communication. Immediately after reception, the data is transferred to an analysis module. Here, the audio and image data are ready to be processed separately.
[0629] Step 3:
[0630] The server uses a generation AI to convert voice data into text. Using speech recognition technology, it extracts keywords and phrases indicating urgency, which serve as the basis for determining the urgency of the report. Simultaneously, it uses an analysis algorithm to analyze the environment and injury status from image data, obtaining information necessary for accurate situation assessment.
[0631] Step 4:
[0632] The server centralizes the analyzed information and prioritizes each report. An AI algorithm automatically prioritizes rescue cases based on the urgency and location of the victims. Furthermore, detailed disaster patterns based on the user's equipment and surrounding environment are simultaneously analyzed.
[0633] Step 5:
[0634] The server promptly notifies emergency services. The notification includes special details about the disaster, location information, and urgency ranking, providing data that enables immediate response. The notification protocol is automatically configured based on a pre-registered database of emergency services.
[0635] Step 6:
[0636] The server continuously monitors and gathers feedback after an alert is received. Based on feedback from emergency services, the system is continuously adjusted, and AI algorithms are improved and databases are updated in real time. Through this process, the overall accuracy and efficiency of the system's response are enhanced.
[0637] (Example 1)
[0638] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0639] During disasters and emergencies, there is a need to ensure that victims can receive emergency assistance quickly and effectively, even when ground-based communication infrastructure is unavailable. Conventional systems have the problem of hindering the transmission of emergency calls due to communication disruptions. Furthermore, there is a challenge in that there are insufficient means to accurately assess the urgency of received information and quickly coordinate with the appropriate emergency services, leading to delays in response.
[0640] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0641] In this invention, the server includes a device that receives emergency calls using satellite communication, a device that converts voice information into text and analyzes image information using generative artificial intelligence, and a device that aggregates the analyzed information and determines the processing order based on the urgency. This makes it possible to quickly determine the urgency and transmit information to the appropriate emergency agency even if ground communication is interrupted.
[0642] "Satellite communication" is a technology that transmits and receives information via communication satellites without using terrestrial communication infrastructure.
[0643] An "emergency call" is communication information sent to request help in emergency situations such as disasters or getting lost.
[0644] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to analyze data and generate information such as speech and text.
[0645] "Converting audio information to text" is the process of analyzing audio and representing it as a corresponding string of characters.
[0646] "Analyzing image information" is a technique that uses image data to identify situations and objects and extract meaningful information.
[0647] "Device" refers to hardware or software designed to perform a series of processes.
[0648] "Urgency" is an indicator that shows the degree to which an emergency response is necessary, and it is a criterion for determining the order in which a rapid response is required.
[0649] "Processing order" refers to guidelines for determining the order in which multiple tasks are performed based on their priority.
[0650] An "emergency agency" is a public or private organization that has the facilities and personnel to respond to emergencies such as disasters and accidents.
[0651] This invention is an emergency notification system that utilizes satellite communication to enable users to receive rapid assistance in emergency situations.
[0652] First, the user uses a dedicated smartphone app. This app is used to send SOS messages in the event of a disaster or emergency. The device is equipped with a satellite communication module and transmits signals via an airborne communication network without using ground-based communication infrastructure. A general-purpose module that supports various protocols and standards could be used as the satellite communication module.
[0653] The server processes data received via satellite communication. Generative artificial intelligence is used to transcribe audio data into text. A general-purpose AI model that excels at speech recognition and can perform effective text conversion is considered. Next, machine learning algorithms are applied to the text data to extract specific keywords indicating urgency. For image data, commonly used image analysis algorithms are used to determine the disaster situation from the video.
[0654] The analyzed data is aggregated as an urgency score, and the processing order is determined based on its importance. This information is notified to the nearest emergency services, and instructions are given to begin immediate action. This process enables rapid and accurate emergency assistance.
[0655] For example, if a user makes a voice call saying, "My house is on fire, please help," the voice data is converted to text by the server, and keywords such as "fire" and "help" are extracted as factors that increase the urgency. As a result, the server processes it as a high-priority call and can immediately transmit the information to fire departments and other relevant organizations.
[0656] Example of a prompt:
[0657] "A user has sent an SOS message from a disaster area. Analyze the audio and video, calculate the urgency, and notify the appropriate agency."
[0658] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0659] Step 1:
[0660] Procedure for users to send SOS messages
[0661] Users launch a dedicated app on their smartphones to record newly occurring emergencies. A voice input function allows them to enter comments describing the situation using voice. They can also capture images and videos of the scene using the camera function and attach them to the SOS message. This data is transmitted from the device to the server via a satellite communication module. Input consists of voice and image data, while output is structured data sent to the server.
[0662] Step 2:
[0663] The server receives the data and converts the audio data into text.
[0664] The server receives the SOS message sent from the terminal. The received audio data is converted into text using a generative AI model. Specifically, the audio waveform data is input to the AI model, which is then analyzed and output as text data.
[0665] Step 3:
[0666] The server extracts urgent keywords from the text.
[0667] The server analyzes the text data and extracts keywords indicating urgency (e.g., "Help," "Fire"). It applies a machine learning algorithm to identify important keywords from a pre-trained model. The input is text data, and the output is the extracted urgency keywords.
[0668] Step 4:
[0669] The process by which the server analyzes image data.
[0670] The server analyzes the received image data to determine the situation and environment of the subject. Using image analysis algorithms, it identifies, for example, collapsed buildings or the condition of people. The input is image data, and the output is the analyzed situational information.
[0671] Step 5:
[0672] A procedure in which the server calculates urgency and determines the processing order.
[0673] The server integrates keywords from voice analysis with image analysis results to calculate the overall urgency. It generates an urgency score using a proprietary calculation model and compares it to other reports based on that score. The output is this score, which is used to determine the processing order.
[0674] Step 6:
[0675] The server notifies the most appropriate emergency agency.
[0676] Based on the determined processing order, the server notifies the most appropriate emergency agency. The information sent to each emergency agency includes the user's location information, transcribed audio information, and analyzed image information. This allows the receiving agency to quickly understand the situation and begin responding. The output is notification data.
[0677] (Application Example 1)
[0678] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0679] In natural disasters and emergencies, there is a need for communication methods that enable rapid and accurate reporting and initial response. However, conventional systems that rely on terrestrial communications struggle to effectively report emergencies when communication infrastructure is destroyed. Furthermore, they lack intuitive operation methods to enable users to report emergencies at the appropriate time. In addition, there are few means to effectively determine the priority of emergency calls, hindering efficient rescue operations.
[0680] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0681] In this invention, the server includes a receiving means for receiving emergency calls using satellite communication, an analysis means for converting voice data into text using a generation AI and analyzing image data, and a prioritizing means for aggregating the analyzed data and determining priorities based on urgency. This realizes an emergency call system that does not depend on communication infrastructure, allows users to make calls intuitively, and enables rapid rescue operations through efficient prioritization.
[0682] "Satellite communication" is a technology that transmits and receives communication signals via artificial satellites, making it possible to transmit signals over a wide area without relying on ground-based communication infrastructure.
[0683] "Generative AI" is an artificial intelligence technology that automatically generates information from data, and has the ability to analyze various types of data such as audio and images to extract text and features.
[0684] "Converting audio data to text" is the process of analyzing audio information and converting it into corresponding textual information.
[0685] "Analyzing image data" refers to the technique of extracting information from images and identifying specific features or patterns.
[0686] "Determining priorities" means evaluating the importance of various data and reports and setting the order in which they require immediate attention.
[0687] "Recording means installed in a portable display device" refers to a function incorporated into a display device that can be easily carried by an individual, which acquires and records audio and video.
[0688] A "startup mechanism" is a function that automatically starts the system when a user performs an action that meets specific conditions.
[0689] A "notification mechanism" is a function for appropriately transmitting analyzed information to relevant authorities and agencies.
[0690] To realize this invention, a portable display device carried by the user, a server, and a satellite communication network are required. The portable display device includes a microphone and camera necessary for acquiring audio and video. When the user gives a voice command in an emergency, the device automatically starts recording audio and video data. This data is transmitted from the device to the server via satellite communication.
[0691] The server uses a generative AI model to convert audio data into text and extracts keywords from that text to assess urgency. Specifically, it uses Google Speech-to-Text to convert audio to text and then analyzes the extracted text with TensorFlow. Additionally, it uses image processing technologies such as OpenCV to evaluate environmental changes and risk levels in video data.
[0692] Based on the analyzed data, the server prioritizes emergency events and generates a list of notifications, ordered from those requiring immediate attention to those requiring it. This enables efficient rescue operations. By utilizing satellite communication, rather than relying on existing communication infrastructure, a highly reliable notification system can be established even in the event of communication disruptions.
[0693] As a concrete example, in a mountainous area where ground communication is cut off due to a natural disaster, if a hiker is about to fall, they can simply say "SOS" into a portable display device, and the appropriate agency will be quickly notified via satellite communication.
[0694] An example of a prompt message for a generative AI model to analyze audio data is: "Design a machine learning model to analyze audio data and extract urgency keywords. Also, propose a method for evaluating the level of risk through image analysis."
[0695] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0696] Step 1:
[0697] The user speaks "SOS" into the portable display device to report an emergency. The device records this audio as sound data and simultaneously acquires video data using its camera. The audio and video inputs are used to make the emergency call.
[0698] Step 2:
[0699] The terminal transmits acquired audio and video data to the server via satellite communication. The input is data recorded within the terminal, which becomes the information to be transmitted to the server. The server receives this data and prepares for data analysis.
[0700] Step 3:
[0701] The server uses a generative AI model to convert speech data into text. It uses Google Speech-to-Text to generate corresponding strings from the speech data and extract urgent keywords. The input is speech data, and the output is textual information. In this process, the user's spoken content is converted into specific textual information.
[0702] Step 4:
[0703] The server uses image analysis techniques such as OpenCV to analyze video data and evaluate the environmental conditions and changes. The input is video data, and the output is feature information of the analyzed video. In this step, high-risk elements within the video are identified.
[0704] Step 5:
[0705] The server integrates the results of voice-to-text and video analysis to determine the priority of notifications based on urgency. This process prioritizes notifications by considering emergency keywords and environmental analysis information. The input is transcribed voice information and video feature information, and the output is prioritized notification data.
[0706] Step 6:
[0707] The server notifies the appropriate emergency services of prioritized alert data. The input is prioritized alert information, which becomes the content of the notification to the emergency services. The server refers to the database, selects the most suitable service based on the content of each piece of information, and sends a notification to encourage a rapid response.
[0708] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0709] This invention relates to an embodiment of an emergency notification system for disasters and emergencies that incorporates an emotion engine to recognize the user's emotions. In this embodiment, the user inputs an emergency message using a dedicated application on their smartphone. The terminal collects emotion data along with voice and image data and transmits it to a server via satellite communication.
[0710] The transmitted data is processed on the server using an analysis method that utilizes generation AI. Within this process, an emotion engine recognizes the user's emotions through voice analysis, and this information is used to determine the level of urgency. For example, based on changes in the user's voice tone and volume, features indicating heightened emotions are extracted and considered as factors that increase the level of urgency.
[0711] Furthermore, image analysis is performed to determine the environment and the extent of injuries from visual information. This information is integrated on a server, and a prioritization mechanism determines the priority of the response. The emotion engine evaluates the impact of emotional information on the severity and urgency of the report by comparing it with past databases.
[0712] Next, the server notifies emergency services of the prioritized information. This process also includes emotional data, enabling emergency services to respond appropriately, taking into account the psychological state of the victims. For example, based on emotional information, emergency services can instruct immediate action on cases that require priority.
[0713] Ultimately, a feedback mechanism improves the accuracy of the operational system through feedback from emergency services. Based on this feedback, the emotion engine algorithm is improved, and the system is continuously adjusted to achieve even greater accuracy in future incidents. This results in a comprehensive emergency call system that enables rapid and highly accurate responses.
[0714] The following describes the processing flow.
[0715] Step 1:
[0716] The user enters an emergency call message and records a voice message through a smartphone application. The device inputs the user's voice into an emotion engine in real time and acquires data to estimate their emotions.
[0717] Step 2:
[0718] The device packages the acquired voice messages, emotion data, and location information and transmits them to the server via satellite communication. The transmitted information also includes the results of the user's emotion analysis.
[0719] Step 3:
[0720] The server inputs the received data into an analysis module and uses a generation AI to convert the voice message into text. Furthermore, it evaluates the user's psychological urgency based on the emotional data provided by the emotion engine.
[0721] Step 4:
[0722] The server analyzes the collected image data using image analysis algorithms to extract information about the damage and the environment. It then references recognized emotion data to comprehensively assess the overall urgency of the situation.
[0723] Step 5:
[0724] The server centralizes the analyzed urgency and sentiment data and uses prioritization mechanisms to determine the priority of rescue operations. Based on this information, cases requiring a rapid response are identified.
[0725] Step 6:
[0726] The server notifies emergency services of priority information. This notification includes detailed disaster information as well as information about the user's emotional state, allowing emergency services to accurately understand the situation and take appropriate action when psychological support is needed.
[0727] Step 7:
[0728] The server continuously monitors feedback from emergency calls and adjusts its emotion engine and analysis methods based on reports from emergency services. This improves the system so that it can provide a more accurate emergency response in the future.
[0729] (Example 2)
[0730] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0731] In the event of a disaster or emergency, accurate analysis of on-site reports is essential for a swift and appropriate response. However, conventional systems lack the accuracy to analyze voice and image data, making it particularly difficult to determine the level of urgency while considering the user's emotional state. Furthermore, coordinating with emergency services presents challenges in responding to the specific needs of the field.
[0732] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0733] In this invention, the server includes a communication device that receives emergency calls using satellite communication, an analysis mechanism that converts voice signals into text data and analyzes image information using generation AI technology, and an emergency assessment device that recognizes the user's emotions and evaluates the urgency by integrating that emotion information with the analyzed data. This makes it possible to accurately determine the priority of responses in the event of a disaster or distress and to provide a swift and appropriate emergency response.
[0734] A "communication device" is a device that uses satellite communication to receive emergency calls from outside.
[0735] An "analysis mechanism" is a device that uses generative AI technology to convert audio signals into text data or to perform detailed analysis of image information.
[0736] An "urgency assessment device" is a device that analyzes the user's emotions, evaluates the local situation along with the integrated data, and determines the level of urgency.
[0737] A "prioritization device" is a device that determines the priority of responses based on the urgency level obtained by an urgency assessment device.
[0738] A "communication device" is a device used to accurately and quickly notify relevant organizations of emergency information.
[0739] A "regulating device" is a device used to improve the analysis mechanism based on feedback from related organizations, thereby enhancing the overall accuracy of the system.
[0740] This invention is a system for effectively making emergency calls during disasters or emergencies. The system consists of a mechanism in which a user inputs an emergency message using a smartphone terminal, and this information is transmitted to a server via satellite communication.
[0741] The device is equipped with the functionality to collect voice and image data. Users use a dedicated smartphone application to make voice notifications and send images. The device utilizes speech recognition technology to convert the user's voice into text data. This conversion uses a generative AI model, providing more advanced speech analysis capabilities.
[0742] The server receives data transmitted from the terminal and processes it using an analysis mechanism. This analysis mechanism, which includes a generative AI, not only converts audio data into text but also has the ability to identify emotions from factors such as tone and volume. Furthermore, it can evaluate the user's visual situation using image data.
[0743] For example, consider a scenario where a user is lost in a mountainous area and uses their device to record a voice message saying, "Please help me, I don't know where I am," and sends it along with an image. An example of a prompt message in this case would be, "We have received your emergency call, will analyze the user's emotions from the voice data, and will develop a rapid response plan."
[0744] The server's urgency assessment device quickly determines the urgency of the situation based on analyzed emotional and visual information. Based on this, the priority determination device organizes the response priorities and notifies those with the highest level of urgency first.
[0745] Finally, the server notifies the relevant organizations of appropriate response information through a communication device. The communication device conveys detailed information about the report, including the user's emotional state, to the relevant organizations, helping them to implement the most appropriate response on the scene.
[0746] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0747] Step 1:
[0748] The user enters an emergency message using a dedicated smartphone application and records audio and image data to the device. This input includes the user's voice and video captured by the smartphone camera. The device formats this data and saves it in an appropriate format for subsequent processing.
[0749] Step 2:
[0750] The device processes recorded audio data and converts it into text data using a generation AI model. The input is the user's audio file, which is then passed through an analysis engine to convert it into text information. This outputs specific textual information. The intonation and tone of the voice are also analyzed simultaneously, and characteristic data indicating emotion is extracted.
[0751] Step 3:
[0752] The device analyzes stored image data and extracts visual features. The input is an image file taken by the user, and an image processing algorithm is used to evaluate the environmental conditions and the state of people. The output includes visual hazards and environmental information derived from the image.
[0753] Step 4:
[0754] The terminal aggregates the converted text data, sentiment data, and image analysis results and sends them to the server via the internet. This transmission process involves data integration and packetization to ensure secure and rapid transmission to the server.
[0755] Step 5:
[0756] The server processes the received data using an analysis mechanism. It analyzes text data to identify urgent elements and evaluates the user's psychological state based on emotional data extracted from the audio. This process efficiently identifies information with high urgency.
[0757] Step 6:
[0758] The server integrates the emotional and image information analyzed by the urgency assessment device to calculate an urgency score. In this step, the urgency of each input data is quantified and prioritized. The final output is a response plan with determined priorities.
[0759] Step 7:
[0760] The server notifies relevant organizations of prioritized information via communication devices. This notification includes the user's location and instructions for responding to high-priority cases. This allows for immediate and appropriate responses on the ground.
[0761] (Application Example 2)
[0762] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0763] Conventional emergency reporting systems often judge the urgency of a situation based on a uniform standard without considering the caller's psychological state or emotions, which can lead to a lack of prompt and appropriate responses tailored to the actual situation. As a result, there is a problem where the prioritization of emergency responses is incorrect, and immediate responses may be delayed. Furthermore, responding without considering the emotional state of the victims can increase their psychological burden.
[0764] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0765] In this invention, the server includes means for receiving emergency calls using satellite communication, means for converting voice information into text using a generation AI and analyzing video information, means for aggregating the analyzed information and determining a ranking based on the degree of urgency, and means for analyzing the user's emotions using an emotion engine and considering this in determining the degree of urgency. This makes it possible to evaluate the overall degree of urgency, including the emotional state of the caller, enabling faster and more appropriate prioritization and response.
[0766] "Satellite communication" is a means of communication that uses artificial satellites to send and receive information between devices located in distant places on Earth.
[0767] An "emergency call receiving device" is a device that quickly receives emergency messages from callers in the event of an abnormal situation or danger.
[0768] "Generative AI" refers to artificial intelligence technology that can analyze various types of data, including audio and video, and extract useful information from them.
[0769] "Auditory information" refers to data that records the words and sounds spoken by a person, and it serves as the basis for analyzing context and emotions.
[0770] "Visual information" refers to visual data acquired through cameras and other imaging devices, and is information that visually captures the state of objects and the environment.
[0771] "Analysis methods" refer to technical techniques for processing collected data and extracting meaningful information from it.
[0772] A "system for determining priority based on urgency" is a system for evaluating the seriousness and urgency of a situation and determining the priority of responses accordingly.
[0773] An "emotion engine" is a technological element that can identify and analyze a person's emotional state from audio and video.
[0774] "The caller's emotional state" refers to the feelings and psychological condition of the person making the emergency call.
[0775] "User" refers to the individual or organization that uses this emergency call system.
[0776] The system that realizes this invention consists of the cooperation of multiple devices and software. When making an emergency call, the user uses a portable information terminal. This terminal is equipped with satellite communication capabilities, enabling highly reliable communication unaffected by the external environment.
[0777] When a user presses the emergency button, the device's microphone records audio and the camera captures video. This information is processed in real time by a generative AI; the audio is converted to text and the video is analyzed. The generative AI uses the Google Cloud Speech-to-Text API to transcribe the audio, while the Google Cloud Vision API is used to analyze the video. This enables detailed analysis of both audio and video.
[0778] The server receives information transmitted from the terminal and analyzes the user's emotions using an emotion engine. In this process, it detects changes in the tone and volume of the user's voice to determine the level of emotional arousal. It also compares this data with past data to determine the degree of urgency. The server sets priorities according to the degree of urgency and immediately notifies emergency services if necessary.
[0779] For example, if a user witnesses someone behaving suspiciously on the street, they can report the emergency situation through the application. Information reflecting the urgency of the voice report is immediately sent to the server, and if action is needed, instructions are quickly given to emergency services. Through this process, it is possible to ensure an appropriate and rapid response in emergencies.
[0780] An example of a prompt to smoothly initiate the analysis of a generative AI model would be: "Based on recent emergency call cases, analyze the emotions in this voice data. Provide guidance to assess the urgency and select the appropriate response route." This allows for more accurate assessment of emotional changes and urgency.
[0781] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0782] Step 1:
[0783] When a user presses the emergency call button on their mobile device, the device begins recording audio and capturing video with its camera. The input consists of the user's audio and video data, which is then converted into a format that can be processed by the AI model. Specifically, this involves collecting the necessary data using the device's microphone and camera.
[0784] Step 2:
[0785] The device uses generative AI to convert recorded audio information into text. This involves analyzing the audio data using the Google Cloud Speech-to-Text API and converting it into text. The input is audio data, and the output is text data. The process includes extracting audio features within the device and processing them as text strings.
[0786] Step 3:
[0787] The device analyzes captured video information using the Google Cloud Vision API. The input is video data, and the output is information derived from the analysis of various visual elements. Specifically, it recognizes objects and environmental characteristics within the image and collects this information as data.
[0788] Step 4:
[0789] The terminal transmits the acquired text information and video analysis results to the server. This is done using satellite communication, ensuring highly reliable data transfer. The input is the transcribed audio information and analysis results, and the output is the information sent to the server.
[0790] Step 5:
[0791] The server analyzes the user's emotions using an emotion engine based on the information it receives. For example, it can infer stress and tension from factors such as voice tone and volume. The input is the user's voice tone information, and the output is the analyzed emotion information. The generation of emotion data within the server is a specific operation.
[0792] Step 6:
[0793] The server aggregates the analysis results and sets priorities based on urgency. During this process, the analysis results are compared with other emergency call data to determine which response should be given the highest priority. The input consists of analyzed emotional information and video data, and the output is the priority ranking. The main operation involves data comparison processing performed internally by the server.
[0794] Step 7:
[0795] Based on the priority determined by the server, a notification is sent to emergency services. This notification includes audio data, video data, and an analyzed level of urgency, enabling an appropriate response. The input is the data of the set priority, and the output is the notification to emergency services. The automatic transmission of data from the server is the specific operation in this processing step.
[0796] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0797] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0798] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0799] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0800] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0801] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0802] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0803] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0804] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0805] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0806] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0807] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0808] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0809] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0810] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0811] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0812] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0813] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0814] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0815] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0816] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0817] The following is further disclosed regarding the embodiments described above.
[0818] (Claim 1)
[0819] A receiving means for receiving emergency calls using satellite communication,
[0820] An analysis method that uses generation AI to convert audio data into text and analyzes image data,
[0821] A prioritization method that aggregates the analyzed data and determines priorities based on urgency,
[0822] Notification methods for notifying emergency agencies,
[0823] A system that includes this.
[0824] (Claim 2)
[0825] The system according to claim 1, characterized in that it analyzes the urgency using speech recognition technology and performs re-analysis as necessary.
[0826] (Claim 3)
[0827] The system according to claim 1, further comprising means for adjusting the analysis means based on feedback from emergency services.
[0828] "Example 1"
[0829] (Claim 1)
[0830] A device that receives emergency calls using satellite communication,
[0831] A device that uses generative artificial intelligence to convert speech information into text and analyze image information,
[0832] A device that aggregates the analyzed information and determines the processing order based on urgency,
[0833] A device that transmits information to emergency services,
[0834] A device that extracts words indicating the degree of urgency from spoken words and determines the extent of the damage through image analysis,
[0835] A system that includes this.
[0836] (Claim 2)
[0837] The system according to claim 1, characterized in that it analyzes the urgency using speech recognition technology and performs re-analysis as necessary.
[0838] (Claim 3)
[0839] The system according to claim 1, further characterized by including a function to adjust the analysis device based on the response from the emergency agency.
[0840] "Application Example 1"
[0841] (Claim 1)
[0842] A receiving means for receiving emergency calls using satellite communication,
[0843] An analysis method that uses generation AI to convert audio data into text and analyzes image data,
[0844] A prioritization method that aggregates the analyzed data and determines priorities based on urgency,
[0845] Recording means mounted on a portable display device that acquires audio and video data at the scene of an emergency,
[0846] An activation mechanism that allows for the notification of an emergency situation by voice and automatically initiates the notification process,
[0847] Notification methods for notifying emergency agencies,
[0848] A system that includes this.
[0849] (Claim 2)
[0850] The system according to claim 1, characterized in that it analyzes the urgency using speech recognition technology and performs re-analysis as necessary.
[0851] (Claim 3)
[0852] The system according to claim 1, further comprising means for adjusting the analysis means based on feedback from emergency services.
[0853] "Example 2 of combining an emotion engine"
[0854] (Claim 1)
[0855] A communication device that receives emergency calls using satellite communication,
[0856] An analysis mechanism that uses generative AI technology to convert audio signals into text data and analyzes image information,
[0857] An urgency assessment device that recognizes the user's emotions and integrates that emotional information with analyzed data to assess the urgency level,
[0858] A priority determination device that determines priorities based on urgency,
[0859] A communication device that notifies relevant organizations of emergency information,
[0860] A system that includes this.
[0861] (Claim 2)
[0862] The system according to claim 1, characterized in that it analyzes the degree of urgency using speech recognition technology and emotion analysis function, and re-analyzes the information as necessary.
[0863] (Claim 3)
[0864] The system according to claim 1, further comprising an adjustment device for adjusting the analysis mechanism based on feedback from related organizations to improve the accuracy of emotion analysis.
[0865] "Application example 2 when combining with an emotional engine"
[0866] (Claim 1)
[0867] A device that receives emergency calls using satellite communication,
[0868] A device that uses generation AI to convert audio information into text and analyzes video information,
[0869] A device that aggregates the analyzed information and determines a ranking based on the degree of urgency,
[0870] A device that notifies emergency services,
[0871] A device that uses an emotion engine to analyze the user's emotions and takes this into consideration when determining the degree of urgency,
[0872] A system that includes this.
[0873] (Claim 2)
[0874] The system according to claim 1, which analyzes the degree of urgency using speech recognition technology and performs re-analysis as necessary.
[0875] (Claim 3)
[0876] The system according to claim 1, further comprising a function to adjust the analysis device based on feedback from an emergency agency. [Explanation of Symbols]
[0877] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A receiving means for receiving emergency calls using satellite communication, An analysis method that uses generation AI to convert audio data into text and analyzes image data, A prioritization method that aggregates the analyzed data and determines priorities based on urgency, Notification methods for notifying emergency agencies, A system that includes this.
2. The system according to claim 1, characterized in that it analyzes the urgency using speech recognition technology and performs re-analysis as necessary.
3. The system according to claim 1, further comprising means for adjusting the analysis means based on feedback from emergency services.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A