System

The system addresses the challenge of rapid emergency response by using image and audio data analysis to automatically contact emergency services and provide first aid, enhancing user response accuracy and survival chances.

JP2026021022APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122704
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Users often struggle to respond quickly and appropriately in emergency situations, failing to accurately grasp the situation and contact the right emergency services, which can lead to decreased chances of survival.

Method used

A system that includes image acquisition, location information, and audio data transmission to a server for analysis, automatically determining the appropriate emergency contact and providing first aid guidance, enabling rapid and accurate responses.

Benefits of technology

Enables users to respond quickly and accurately in emergencies, improving survival rates by automatically contacting emergency services and providing first aid guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021022000001_ABST
    Figure 2026021022000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for activating image acquisition means and acquiring location information using location information acquisition means when triggered by a user; transmission means for transmitting acquired image data, audio data, and location information; analysis means for analyzing data transmitted by the transmission means and identifying an emergency condition; contact means for determining an appropriate emergency contact and automatically initiating a call based on an emergency condition identified by the analysis means; and guidance means for providing first aid guidance to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When encountering an emergency, users often lose their composure and find it difficult to respond quickly and appropriately. Specifically, it is often difficult for users to accurately grasp the situation and contact the appropriate emergency service. Furthermore, they may not be able to obtain appropriate instructions for first aid. In such situations, the chances of survival are likely to decrease. The present invention aims to solve these problems and provide a system that supports quick and appropriate responses in emergency situations. [Means for solving the problem]

[0005] In order to solve the above problems, the present invention provides a system having the following means. The system includes means for activating an image acquisition means when triggered by a user and acquiring location information using a location information acquisition means, and a transmission means for transmitting the acquired image data, audio data, and location information. The system further includes an analysis means for analyzing the transmitted data and identifying an emergency situation. The system also includes a contact means for determining the most appropriate emergency contact based on the analyzed information and automatically initiating a call. In addition, the system includes a guidance means for providing the user with first aid guidance. This configuration enables users to respond quickly and accurately in an emergency, which is expected to improve the survival rate.

[0006] "User" refers to a person who uses the system to deal with an emergency.

[0007] "Image acquisition means" refers to the function of collecting visual data such as videos and photographs using a camera.

[0008] "Location Acquisition Method" refers to the ability to determine your current geographic location using GPS or other location services.

[0009] "Transmission means" refers to a function for transmitting collected data to a remote server.

[0010] "Analysis means" refers to the data processing functions used to identify and classify emergencies based on the transmitted data.

[0011] An "emergency situation" refers to a dangerous situation or accident that requires a prompt response.

[0012] "Contact" refers to a calling feature for contacting appropriate emergency services based on an identified emergency condition.

[0013] "Guidance means" refers to a function for guiding the user to first aid or other appropriate responses.

[0014] "Acoustic analysis means" refers to a function that analyzes the surrounding sound environment based on audio data.

[0015] A "language model" refers to a data model that has been pre-trained for natural language processing. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is a system for responding quickly and appropriately to emergencies. The system is built around a smartphone application that allows users who encounter an emergency to automatically contact the appropriate emergency services with the push of a button.

[0038] System configuration

[0039] This system consists of a user's smartphone device (hereinafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. A multimodal generation AI runs on the server side and analyzes the transmitted data.

[0040] System Operation

[0041] 1. Pressing the emergency button and collecting data

[0042] When a user presses the emergency button, the device activates the camera and starts recording video from that point onward, simultaneously recording audio data using the microphone, and obtaining the device's current location using the GPS module.

[0043] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[0044] 2. Data submission and analysis

[0045] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[0046] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[0047] 3. Determine emergency services and initiate a call

[0048] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[0049] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[0050] 4. First Aid and Navigation

[0051] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station, allowing the device to provide turn-by-turn navigation to the user.

[0052] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[0053] This system will enable users to respond quickly and accurately in emergency situations, which is expected to improve the chances of survival. In addition, by utilizing speech analysis methods and language models that work even in unstable communication environments, appropriate support can be provided in any situation.

[0054] The processing flow will be explained below.

[0055] Step 1:

[0056] When a user presses the emergency button, the device immediately activates the camera and starts recording video. At the same time, the microphone is activated to record surrounding audio. The device also uses the GPS module to obtain its current location.

[0057] Step 2:

[0058] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[0059] Step 3:

[0060] The server adds the received data packets to the analysis queue and activates the multimodal generation AI. The server analyzes the video data to identify emergency situations (e.g., traffic accidents, fires, collapsed people, etc.) and analyzes the surrounding sound environment from the audio data.

[0061] Step 4:

[0062] Based on the analysis results, the server determines the type and severity of the emergency, identifies the current geographic location from the location information, and lists the nearest emergency services (e.g., police station, fire station, hospital, etc.).

[0063] Step 5:

[0064] The server selects the most appropriate emergency service contact based on the determined emergency situation and returns that information to the device.

[0065] Step 6:

[0066] The device responds to the received instructions and automatically calls the selected emergency service contact, and plays a voice message generated by the server explaining the situation.

[0067] Step 7:

[0068] After the call is made, the device will begin providing first aid instructions to the user. The server will calculate the optimal route to the nearest hospital or fire station based on the user's location and send that information to the device.

[0069] Step 8:

[0070] The device will display the calculated optimal route information to the user, initiate turn-by-turn navigation, and even if the connection is unstable, it can continue to provide first aid instructions using pre-installed language models.

[0071] Step 9:

[0072] By following the first aid and navigation instructions, users can facilitate appropriate responses at the scene until emergency services arrive, which can lead to faster and more accurate responses and potentially improved survival rates.

[0073] Example 1

[0074] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0075] When encountering an emergency, it can be difficult for individuals to contact emergency services quickly and appropriately. This is especially true when users are in a state of panic or have difficulty explaining the situation, which can delay appropriate responses. Furthermore, there is a risk that damage may escalate due to users not knowing how to respond appropriately. There is a need for a system that can resolve these issues and provide a fast and appropriate emergency response.

[0076] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0077] In this invention, the server includes means for activating an image acquisition means and a voice acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, voice data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for calculating an optimal route based on the user's location information and providing navigation, thereby enabling the user to respond to an emergency quickly and accurately.

[0078] "Image acquisition means" refers to a device or function that uses a camera or other image sensor to acquire video or image information about the surroundings.

[0079] An "audio capture means" is a device or function that uses a microphone or other audio sensor to record or capture ambient sounds.

[0080] "Location Information Acquisition Means" means a device or function that acquires current geographic location information using GPS or other location determination technology.

[0081] The "transmission means" is a function that assembles the acquired data into a single packet and transmits it via the Internet or other communication means.

[0082] "Analysis means" refers to software or hardware that analyzes received data and identifies specific conditions or information.

[0083] "Multimodal generation AI" is an artificial intelligence technology that integrates and analyzes multiple modals (input formats) such as images, audio, and text.

[0084] "Emergency Contacts" are contact details for agencies and services (e.g., ambulance, police, fire department, etc.) that will best respond to a particular emergency.

[0085] "Contact methods" are devices or functions that automatically initiate a call to an emergency contact and convey the necessary information.

[0086] The "guidance means" is a function for providing first aid and navigation information to the user.

[0087] "Navigation means" is a function that calculates the optimal route based on the user's location information and provides real-time directions.

[0088] MODE FOR CARRYING OUT THE INVENTION

[0089] The present invention is a system for responding promptly and appropriately to an emergency situation, and functions in cooperation with a user, a terminal, and a server. A specific embodiment of this system will be described below.

[0090] 1. Components and Hardware Configuration

[0091] Terminal

[0092] The device is a smartphone owned by the user and is equipped with the following hardware:

[0093] Camera: Captures video and images of the surrounding area.

[0094] Microphone: Records surrounding sounds.

[0095] GPS module: Obtains current location information.

[0096] These hardware components function as an "image acquisition means," "audio acquisition means," and "location information acquisition means." In addition, a dedicated application is installed, and these functions are automatically activated when the emergency button is pressed.

[0097] server

[0098] The server will be deployed in a cloud environment and will utilize the following software and technologies:

[0099] Multimodal generative AI: Integrated analysis of video, audio, and location data to identify emergencies.

[0100] Communication means: Receives data sent from the terminal and sends the analysis results to the terminal.

[0101] 2. Usage and Data Processing

[0102] Pressing the emergency button

[0103] When a user presses the emergency button within a smartphone application, the device automatically activates the camera, microphone, and GPS to collect data.

[0104] Sending data

[0105] The terminal combines the collected video data, audio data, and location information into a single data packet, which is then transmitted to a server via the Internet by a transmitting means.

[0106] Data analysis

[0107] The server analyzes the received data packets using an analytical method (multimodal generative AI), which identifies the type of emergency and selects the appropriate emergency services (ambulance, police, fire department, etc.).

[0108] Example: If the server analyzes the data it receives and determines that it is a "traffic accident" based on video and audio, the analysis results will also include the extent of injuries and the danger level of the scene.

[0109] Starting a call

[0110] The device receives instructions from the server and automatically calls the designated emergency service, explaining the situation using a voice message generated by the server.

[0111] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, providing details of the emergency situation via voice.

[0112] First Aid Guide and Navigation

[0113] While waiting for the ambulance to arrive, the device provides a pre-prepared first aid guide from the server. The server also tracks the user's location in real time and calculates the optimal route to the nearest hospital or fire station. As a navigation tool, the device provides turn-by-turn navigation information to the user.

[0114] Example: The device provides voice and text instructions for CPR and displays directions to the nearest hospital.

[0115] Prompt Sentence Examples

[0116] An example of a prompt to be input to the generative AI model when a user encounters an emergency:

[0117] "I was involved in a traffic accident and someone was injured. My current location is latitude xx.xxxx, longitude yy.yyyy."

[0118] "A fire has broken out and there are many injured people nearby. Your current location is latitude xx.xxxx, longitude yy.yyyy."

[0119] In this way, the entire system works together to provide comprehensive support in emergencies, enabling users to quickly take the most appropriate action.

[0120] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0121] Step 1:

[0122] User operations

[0123] The user presses the emergency button in the application.

[0124] Input: User touch input.

[0125] Output: Emergency button pressed event.

[0126] Specific action: The user taps the emergency button displayed on the smartphone screen.

[0127] Step 2:

[0128] Device behavior: Data collection

[0129] The device detects when the emergency button is pressed and activates the camera, microphone, and GPS module to collect data.

[0130] Input: Emergency button press event.

[0131] Output: Video data, audio data, location information.

[0132] Specific behavior:

[0133] The camera will activate and record video of the emergency.

[0134] The microphone records the audio.

[0135] The GPS module obtains the current location information.

[0136] Step 3:

[0137] Device operation: Data transmission

[0138] The device bundles the collected data (video data, audio data, location information) into a single data packet and sends it to the server.

[0139] Input: Video data, audio data, location information.

[0140] Output: Data packets sent to the server.

[0141] Specific behavior:

[0142] Video, audio, and location information are packaged into packets.

[0143] The packet is sent over the Internet to a server.

[0144] Step 4:

[0145] Server Operation: Data Analysis

[0146] The server analyzes the received data packets and runs a multimodal generative AI to identify emergencies.

[0147] Input: Data packet.

[0148] Output: Emergency type, detailed analysis results.

[0149] Specific behavior:

[0150] The received video data, audio data, and location information are expanded.

[0151] Multimodal generative AI analyzes the data, identifies the type of emergency (e.g., traffic accident, fire, sudden illness, etc.), and determines the detailed situation.

[0152] Step 5:

[0153] Server Behavior: Emergency Services Decision

[0154] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, or fire department).

[0155] Input: Emergency type, detailed analysis results.

[0156] Output: Emergency services contact information.

[0157] Specific behavior:

[0158] Retrieve the best contact information from a database for the type of emergency.

[0159] Send contact information for selected emergency services to the device.

[0160] Step 6:

[0161] Device Action: Initiating a call

[0162] The device receives instructions from the server and automatically calls emergency services. During the call, it plays a voice message generated by the server explaining the situation.

[0163] Input: Emergency services contact information, server-generated voice message.

[0164] Output: Calls to emergency services, playback of voice messages.

[0165] Specific behavior:

[0166] Your device will automatically dial the emergency services number.

[0167] Once the call is connected, a server-generated voice message is played, detailing the emergency.

[0168] Step 7:

[0169] Device Operation: First Aid and Navigation

[0170] While waiting for the ambulance to arrive, the device will provide the user with first aid guidance and navigate the optimal route to the nearest hospital or fire station.

[0171] Input: First aid guide information provided by the server, real-time location information.

[0172] Output: First aid guide, navigation information.

[0173] Specific behavior:

[0174] The device provides the user with pre-programmed first aid instructions (audio, text, video).

[0175] The server calculates the optimal route based on the user's location information and sends navigation information to the terminal.

[0176] The device displays real-time turn-by-turn navigation to the user.

[0177] (Application example 1)

[0178] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0179] Conventional emergency response systems have the drawback of making it difficult for users to respond quickly and appropriately when they encounter an emergency. In particular, they require complex procedures, such as initiating a call to an emergency contact and providing first aid instructions, which often prevents users from responding calmly. Furthermore, they lack the functionality to provide an optimal route based on the user's current location, making it difficult to evacuate to a safe location efficiently. Therefore, there is a need for a system that supports users' quick and accurate actions in emergency response.

[0180] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0181] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for the guidance means to calculate an optimal travel route based on the user's current location and provide navigation, thereby enabling the user to respond quickly and accurately to an emergency situation and efficiently evacuate to a safe place.

[0182] - A "trigger" is an input device such as a button or switch that a user operates to initiate a specific action.

[0183] "Image acquisition means" refers to a device that collects video data using a camera or the like.

[0184] "Location information acquisition means" is a device that identifies the current location using a GPS module or the like.

[0185] "Audio data" refers to a recording of sound captured using a microphone or the like.

[0186] "Transmission means" refers to a device or process that transmits collected data to a server via the Internet or other communication means.

[0187] "Analysis means" refers to algorithms or software that process the received data and identify emergency situations.

[0188] A "call" is a means of making a voice communication, which may be made over a telephone line or the Internet.

[0189] "Contact Assistance" means any device or software that automatically places a call to the appropriate emergency service.

[0190] A "guiding means" is a device or software that provides useful information to a user.

[0191] A "First Aid Guide" is information that provides instructions on how to provide basic medical assistance in an emergency.

[0192] "Navigation" is a route guidance function that guides a user to a specific destination.

[0193] "Current location of the user" is real-time location information of the user identified by the location information acquisition means.

[0194] This invention is a system for responding quickly and accurately to emergencies, and is built around a user's smartphone terminal. The configuration and operation of the system are described in detail below.

[0195] 1. System Configuration

[0196] This system consists of a user's smartphone terminal (hereinafter referred to as the terminal), a server connected via the Internet, and an emergency service. The terminal is equipped with a camera, microphone, GPS, and the application of this invention. A multimodal generation AI runs on the server side.

[0197] 2. System Operation

[0198] Pressing the emergency button and collecting data

[0199] When the user presses the emergency button, the application installed on the device activates the camera and starts recording video from that point. At the same time, it also records audio data using the microphone. It also obtains the current location information using the GPS module.

[0200] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[0201] Data transmission and analysis

[0202] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[0203] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[0204] Determining emergency services and initiating a call

[0205] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[0206] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[0207] First Aid and Navigation

[0208] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station based on the user's current location, allowing the device to provide turn-by-turn navigation to the user.

[0209] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[0210] Prompt Sentence Examples

[0211] This is an emergency. A user has suddenly lost consciousness at home. Location, video, and audio data are attached below. Please take appropriate action.

[0212] Location: Tokyo, Japan

[0213] Video data: emergency_video.avi

[0214] Audio data: emergency_audio.wav

[0215] Hardware and software used

[0216] Device: Smartphone (camera, microphone, GPS)

[0217] Server: Internet Server

[0218] Software: Multimodal generative AI, Python, OpenCV, PyAudio, Geopy, Requests library

[0219] This allows users to respond quickly and accurately to emergencies and efficiently evacuate to a safe location. In addition, the server tracks users' locations in real time and provides first aid guidance and optimal route navigation, which is expected to improve survival rates.

[0220] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0221] Step 1:

[0222] The device detects the pressing of the emergency button. The input is the user pressing the button, and the output is the activation of the camera, microphone, and GPS. When the user performs a trigger operation, the device detects this event and activates the camera, microphone, and GPS sensors.

[0223] Step 2:

[0224] The device records video with the camera and audio with the microphone. The input is data information from the camera and microphone, and the output is a video file and an audio file. The device starts the camera, records video for a certain period of time (e.g., 10 seconds), and simultaneously records audio data using the microphone. These data are then saved to a file.

[0225] Step 3:

[0226] The device uses GPS to obtain current location information. The input is the signal from the GPS module and the output is location data. The device uses the GPS module to obtain real-time location information (latitude and longitude) and saves the data in the application.

[0227] Step 4:

[0228] The device collects video data, audio data, and location information and combines them into packets. The input is each file and location data, and the output is a data packet. The device combines the video file, audio file, and location information into a single packet.

[0229] Step 5:

[0230] The terminal sends a data packet to the server. The input is the data packet, and the output is a notification of successful data transmission to the server. The terminal sends the data packet to the server via the Internet and confirms its success.

[0231] Step 6:

[0232] The server analyzes the received data. The input is a data packet, and the output is the emergency identification result. The server separates video, audio, and location information from the received data packet, analyzes this data using a generative AI model, and identifies the type of emergency.

[0233] Step 7:

[0234] The server determines the most appropriate emergency contact based on the analysis results. The input is the analysis results, and the output is emergency contact information. Based on the analysis results, the server selects the most appropriate emergency service (ambulance, police, fire department, etc.).

[0235] Step 8:

[0236] The device receives instructions from the server and automatically calls the appropriate emergency service. The input is emergency contact information, and the output is a successful call. The device automatically calls the emergency contact provided by the server and starts the call.

[0237] Step 9:

[0238] The device automatically explains the situation using a voice message generated by the server. The input is the voice message from the server, and the output is the completion of sending the voice message. The device plays the voice message generated by the server and explains the situation to the emergency services.

[0239] Step 10:

[0240] The terminal provides the user with a first aid guide prepared in advance on the server side. The input is first aid guide information, and the output is completion of guide provision. The terminal provides the first aid guide provided by the server to the user in voice or text format.

[0241] Step 11:

[0242] The server tracks the user's real-time location and calculates the optimal route to the nearest hospital or fire station. The input is the user's current location information and the output is the optimal route information. The server uses GPS data to identify the user's current location and calculates the optimal route to the nearest emergency facility.

[0243] Step 12:

[0244] The terminal provides optimal route navigation to the user. The input is optimal route information, and the output is navigation completion. The terminal provides turn-by-turn navigation to the user based on the optimal route information provided by the server.

[0245] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0246] The present invention is a system that takes into consideration the user's emotions in an emergency and can respond quickly and appropriately. This system provides a means for responding to an emergency through a smartphone app, and provides appropriate guidance according to the user's stress and impatience.

[0247] System configuration

[0248] This system consists of the user's smartphone device (hereafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. The server runs a multimodal generation AI and an emotion recognition engine, which analyzes the transmitted data.

[0249] System Operation

[0250] 1. Pressing the emergency button and collecting data

[0251] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously enables the microphone to record audio, acquires location information using the GPS module, and packetizes the audio data for transmission to the server.

[0252] Example: When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and obtaining location information.

[0253] 2. Data submission and analysis

[0254] The device then combines the captured video and audio files, along with location information, into a single data packet and sends it to the server, which then queues the data for analysis and activates the multimodal generation AI and emotion recognition engine.

[0255] Example: The server analyzes the data it receives, identifies traffic accidents and emergencies from voices, and recognizes emotions from the user's tone of voice.

[0256] 3. Emotion Recognition and Emergency Service Decisions

[0257] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as anxiety, fear, panic, etc. The server then uses the analysis results to determine the type and severity of the emergency and selects the appropriate emergency services (e.g., police, ambulance, fire department, etc.).

[0258] Example: If the server detects that the user is panicking, it will determine the level of urgency as high and send instructions to the device to immediately contact an ambulance and the police.

[0259] 4. Automatic contact and first aid provision

[0260] The device receives instructions from the server and automatically calls the selected emergency service contact. The device plays a voice message generated by the server, explaining the situation. Furthermore, the device adjusts the first aid guidance content according to the user's emotional state and provides it to the user.

[0261] Example: For a user in a panicked state, a calming voice will provide instructions on CPR and how to treat injuries.

[0262] 5. Providing location information and navigation

[0263] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information, and the device displays this information to the user and provides turn-by-turn navigation.

[0264] Examples include providing detailed directions to the nearest hospital to help users avoid getting lost, and continuing to provide first aid instructions even when communication is unstable, using pre-installed language models.

[0265] 6. Providing Additional Information

[0266] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions, so that emergency services can make appropriate preparations before arriving on scene.

[0267] Example: If the user is very disoriented, this information can also be communicated to emergency responders so that they can provide psychological support when they arrive on scene.

[0268] In this way, the system, including the emotion recognition engine, enables a fast and accurate response in emergency situations, which is expected to increase the user's sense of security and improve the overall survival rate.

[0269] The processing flow will be explained below.

[0270] Step 1:

[0271] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously activates the microphone to capture surrounding audio, and uses the GPS module to obtain the device's current location.

[0272] Step 2:

[0273] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[0274] Step 3:

[0275] The server adds the received data packets to an analysis queue and activates a multimodal generation AI and emotion recognition engine, which analyzes video data to identify emergency situations and recognizes user emotions from audio data.

[0276] Step 4:

[0277] The server's emotion recognition engine analyzes the user's tone of voice, speed, and word choice from the audio data to determine whether the user is anxious, calm, or panicked.

[0278] Step 5:

[0279] Based on the analysis results, the server determines the type and severity of the emergency, selects the most appropriate emergency service contact, and sends this information back to the device.

[0280] Step 6:

[0281] Based on the instructions received, the device automatically calls the selected emergency service contact and plays a server-generated voice message explaining the situation.

[0282] Step 7:

[0283] The device provides first aid instructions to the user. If the server's emotion recognition engine determines that the user's emotional state is impatient or panicked, the device adjusts the content and speed of the instructions to help the user calm down.

[0284] Step 8:

[0285] The server calculates the optimal route to the nearest hospital or fire station based on real-time location information and sends that information to the device, which then displays the calculated route information to the user and begins turn-by-turn navigation.

[0286] Step 9:

[0287] The emotion recognition engine takes into account the user's emotions and provides additional information to emergency contacts, allowing them to make appropriate preparations before emergency services arrive on scene.

[0288] Step 10:

[0289] By following the instructions on the device and receiving first aid and navigation, users can take appropriate action until emergency services arrive, which will enable quick and accurate response and is expected to improve the survival rate.

[0290] Example 2

[0291] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0292] Conventional emergency response systems are required to respond not only based on user reports but also by taking into account the user's emotional state. However, current systems have difficulty accurately recognizing the user's emotions and providing appropriate guidance and emergency measures. Another challenge is providing accurate responses in environments with unstable communication conditions.

[0293] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0294] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for adjusting first aid guidance according to the user's emotional state, and a means for calculating an optimal route based on the user's real-time location information and providing navigation, thereby enabling a quick and accurate emergency response that also takes the user's emotional state into consideration.

[0295] 1. "Image capture means" refers to a device that activates a camera and captures video or still images when a user presses a trigger.

[0296] 2. "Location information acquisition means" refers to a device that acquires location information of the user's current location using a GPS module or the like.

[0297] 3. "Transmitting means" refers to a device that packetizes the acquired image data, audio data, and location information and transmits them to the server.

[0298] 4. "Analysis means" is a device that analyzes the data transmitted by the transmission means and identifies an emergency situation or the user's emotions.

[0299] 5. "Contact Means" means a device that determines the appropriate emergency contact based on the emergency situation identified by the Analysis Means and automatically initiates a call.

[0300] 6. "Guidance means" is a device that provides the user with first aid guidance and appropriate actions according to their emotional state.

[0301] 7. "Means for providing navigation" means a device that calculates the optimal route based on the user's real-time location information and provides turn-by-turn navigation.

[0302] 8. "Acoustic analysis means" refers to a device that analyzes the surrounding acoustic environment and the user's emotions based on audio data.

[0303] 9. "Language model" refers to a model that is installed to function even in environments with unstable communication and is used to generate emergency and guidance messages.

[0304] MODE FOR CARRYING OUT THE INVENTION

[0305] This invention is an emergency response system that takes into account the user's emotions and can respond quickly and appropriately in an emergency. This system operates via a smartphone app and handles emergencies using an emotion recognition engine and multimodal generation AI.

[0306] Hardware Configuration

[0307] The hardware configuration of this system is as follows:

[0308] 1. Device: A smartphone carried by the user, equipped with the following functions:

[0309] Camera features

[0310] Microphone function

[0311] GPS Modules

[0312] Internet communication function

[0313] 2. Server: A back-end system connected to terminals via the Internet, and includes the following components:

[0314] Multimodal generation AI (audio data analysis, image data analysis)

[0315] Emotion Recognition Engine

[0316] Data Analysis Engine

[0317] Software Configuration

[0318] 1. Terminal application: A smartphone application used by the user that performs the following actions in the event of an emergency.

[0319] Emergency button activates camera and microphone

[0320] Obtaining location information

[0321] Packetizing acquired data (video, audio, location information) and sending it to the server

[0322] 2. Server application: Software that runs on the server side and provides the following functions:

[0323] Analyzing received data

[0324] Recognizing the user's emotional state

[0325] Determining the type and severity of an emergency

[0326] Selection of emergency contacts and automatic call instructions

[0327] Generate and submit a first aid guide

[0328] Real-time location-based navigation

[0329] Example of operation

[0330] To give a concrete example of how this works, consider the following scenario:

[0331] 1. If the emergency button is pressed:

[0332] A user encounters a traffic accident and presses the emergency button on their smartphone.

[0333] The device immediately activates the camera and records video of the accident scene.

[0334] At the same time, the microphone is enabled to record audio and location information is obtained using GPS.

[0335] The terminal transmits this data to the server.

[0336] 2. Data transmission and analysis:

[0337] The server adds the received video, audio, and location information to an analysis queue and activates the multimodal generation AI and emotion recognition engine.

[0338] The server analyzes the user's emotional state (e.g., panic) from the tone of the voice and determines the appropriate emergency response.

[0339] 3. Contacting the appropriate emergency services:

[0340] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and issues an automatic call instruction.

[0341] The device automatically contacts emergency services and plays a server-generated voice message explaining the situation.

[0342] Prompt Sentence Examples

[0343] Examples of prompts include:

[0344] "A user encounters a traffic accident and presses the emergency button. The mobile device collects video, audio, and GPS data and sends it to a server. The server analyzes the urgency of the accident and coordinates the appropriate emergency services. Please explain how the system works."

[0345] This system enables a fast and effective response in emergency situations while taking into consideration the user's feelings, which is expected to improve the user's sense of security and increase the overall survival rate.

[0346] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0347] Step 1:

[0348] The user presses the emergency button.

[0349] Input: The user presses the emergency button.

[0350] Output: The emergency button press signal is transmitted to the terminal.

[0351] The device immediately activates the camera and starts recording video, while simultaneously activating the microphone for audio recording and the GPS module for location information.

[0352] Specific operation: Video recording, audio recording, and location information acquisition of the accident scene are carried out simultaneously.

[0353] Step 2:

[0354] The terminal combines the acquired video file, audio file, and location information into a single data packet and sends it to the server.

[0355] Input: Video data acquired by the camera, audio data recorded by the microphone, and location information acquired by GPS.

[0356] Output: A data packet is generated and sent to the server.

[0357] Specific operation: Various data is packetized and sent to a server via the Internet.

[0358] Step 3:

[0359] The server adds the received data to a queue for analysis and processing, and activates the multimodal generation AI and emotion recognition engine.

[0360] Input: Data packets sent from the terminal.

[0361] Output: The analysis results include the type of emergency and the user's emotional state.

[0362] Specific operation: The received data is added to the analysis queue, and the AI ​​and emotion recognition engine are activated to begin data analysis.

[0363] Step 4:

[0364] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as impatience, fear, or panic.

[0365] Input: Audio data.

[0366] Output: Analysis results showing the user's emotional state.

[0367] Specific operation: Emotion analysis is performed using voice data to determine the user's emotional state.

[0368] Step 5:

[0369] Based on the analysis results, the server determines the type and urgency of the emergency and selects the appropriate emergency service (e.g., police, ambulance, fire department, etc.).

[0370] Input: The type of emergency situation as the analysis result and the user's emotional state.

[0371] Output: Instructions for optimal emergency service selection.

[0372] Specific operation: Based on the analysis results, the level of urgency is determined to be high, and appropriate emergency services are immediately selected.

[0373] Step 6:

[0374] The device receives instructions from the server and automatically calls the selected emergency service contact.

[0375] Input: Instructions from the server.

[0376] Output: Initiation of automatic contact with emergency services.

[0377] Specific operation: The device makes an automatic call and plays a voice message generated by the server to explain the situation.

[0378] Step 7:

[0379] The device provides first aid guidance according to the user's emotional state.

[0380] Input: Analysis results indicating the user's emotional state.

[0381] Output: First aid guide depending on emotional state.

[0382] Specific actions: Provide breathing techniques and first aid instructions in a calm voice.

[0383] Step 8:

[0384] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information.

[0385] Input: The user's real-time location.

[0386] Output: Optimal route information.

[0387] Specific operation: Analyze real-time location information and calculate the optimal route.

[0388] Step 9:

[0389] The device displays optimal route information to the user and provides turn-by-turn navigation.

[0390] Input: Optimal route information.

[0391] Output: Navigation guide.

[0392] Specific behavior: Support the user by providing detailed directions to the nearest hospital.

[0393] Step 10:

[0394] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions.

[0395] Input: Data indicating the user's emotional state.

[0396] Output: Additional information for emergency contacts.

[0397] Specific behavior: If the user is confused, this information will also be conveyed to emergency responders so that psychological support can be prepared when they arrive at the scene.

[0398] (Application example 2)

[0399] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0400] In the event of an emergency, if a user feels anxious or scared, it can be difficult to respond appropriately. Furthermore, conventional emergency response systems do not take the user's emotional state into account, which can lead to inadequate responses. Furthermore, the system must function even in environments with unstable communications. The present invention aims to solve these problems and provide a system that allows users to respond to emergencies with peace of mind.

[0401] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0402] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, an emotion recognition means for analyzing the user's emotional state identified by the analysis means, a contact means for determining an appropriate emergency contact based on the analysis result of the emotion recognition means and automatically initiating a call, and a guidance means for providing guidance on appropriate first aid according to the user's emotional state. This enables a quick and appropriate emergency response while taking the user's emotional state into consideration.

[0403] "Image capture means" refers to a technical means for capturing an image or video using a device such as a camera when a user activates a trigger.

[0404] "Location information acquisition means" refers to a technical means for determining the current location of a terminal using GPS or other positioning systems.

[0405] The "transmission means" is a means having a function of transmitting the acquired image data, audio data, and location information to a server or other external system.

[0406] "Analysis Means" means the technical means used to identify an emergency situation and determine appropriate emergency contacts based on the Transmitted Data.

[0407] "Emotion recognition means" is a technical means for analyzing the user's emotional state from voice data, etc., and identifying emotions such as impatience or fear.

[0408] A "contact method" is a method that has the ability to automatically initiate a call to an emergency contact based on an identified emergency condition and emotional state.

[0409] The "guidance means" is a technical means for providing a first aid guide in response to an emergency situation faced by a user, taking into consideration the user's emotional state.

[0410] MODE FOR CARRYING OUT THE INVENTION

[0411] To implement this invention, the user's smartphone terminal, the server, and the emergency service work together. The hardware and software used in each step and the processing content thereof will be specifically described below.

[0412] System configuration

[0413] 1. Hardware

[0414] Device (smartphone): Smartphones are equipped with a camera, microphone, and GPS module. These functions are used to acquire images, audio, and location information.

[0415] Server: The server runs a multimodal generation AI and emotion recognition engine, which analyzes the transmitted data.

[0416] 2. Software

[0417] Image capture tool: Software that activates the camera and captures images and videos.

[0418] Audio recording software: Records audio using a microphone and generates audio data.

[0419] Location information acquisition software: Uses the GPS module to determine the device's current location.

[0420] Data transmission software: Collects the acquired data into a single data packet and sends it to a server over the Internet.

[0421] Emotion recognition engine: Analyzes the user's emotional state from voice data and identifies emotions such as impatience and fear. Specifically, it uses an emotion recognition model from the Transformers library.

[0422] Multimodal generative AI: Identifies emergencies based on submitted data and generates appropriate responses. Uses the nlptown / bert-base-multilingual-uncased-sentiment model.

[0423] Contact Method: Automatically call emergency contacts based on the analyzed results.

[0424] Guidance: Provides first aid guidance according to the user's emotional state. Even if communication is unstable, the guidance continues using a pre-installed language model.

[0425] System Operation

[0426] 1. Pressing the emergency button and collecting data

[0427] When a user presses the emergency button, the device's camera automatically activates and begins recording images and videos. At the same time, the microphone activates and records audio. The GPS module also activates and collects location information. This data is packetized within the device and sent to the server.

[0428] 2. Data Analysis and Emotion Recognition

[0429] The server analyzes the received data and identifies the user's emotional state from the user's voice tone and environmental sounds. An emotion recognition engine analyzes the voice data and identifies whether the user is expressing emotions such as anxiety or fear, and determines the appropriate emergency measures.

[0430] Examples:

[0431] When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and acquiring location information. The server analyzes this data to understand the circumstances of the accident and the user's emotional state.

[0432] 3. Selecting and contacting emergency services

[0433] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and automatically calls the contacts. The server generates a voice message explaining the situation.

[0434] 4. First aid information

[0435] The system takes into account the user's emotional state and provides the most appropriate first aid guidance. For example, if a user is in a panic, a slow voice will guide them through CPR and how to treat injuries.

[0436] Examples:

[0437] By inputting the following prompts into the generative AI model, specific emergency response instructions can be generated.

[0438] Example prompt:

[0439] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[0440] As a result, it is possible to take prompt and appropriate emergency measures while taking into consideration the emotional state of the user.

[0441] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0442] Step 1:

[0443] When a user presses the emergency button, the device immediately activates the camera and starts recording images and videos. At the same time, it activates the microphone to record audio and uses the GPS module to obtain location information. These data are packetized and ready to be transmitted. The input is the user pressing the button, and the output is image data, audio data, and location information.

[0444] Step 2:

[0445] The device collects the acquired image data, audio data, and location information into a single data packet and sends it to a server via the Internet. The input is the image data, audio data, and location information, and the output is the data packet sent to the server.

[0446] Step 3:

[0447] The server analyzes the received data packets using a multimodal generation AI and emotion recognition engine. First, it identifies the emergency situation from image data and location information, and then analyzes the audio data to recognize the user's emotional state. The input is the data packets, and the output is the emergency situation identification result and the user's emotional state recognition result.

[0448] Step 4:

[0449] The server analyzes the voice data using an emotion recognition engine to determine whether the user is expressing emotions such as anxiety or fear. For example, it uses an emotion recognition model from the transformers library. The input is the voice data, and the output is the identification of the emotional state.

[0450] Step 5:

[0451] The server determines the appropriate emergency contact based on the analysis results and instructs the device to automatically initiate a call. The server generates the necessary call content using a generative AI model and sends it to the device. The input is the emergency state and emotional state identification results, and the output is the emergency contact and call content.

[0452] Step 6:

[0453] The device receives instructions from the server and automatically calls the appropriate emergency contact. It responds by playing a voice message generated by the server and explaining the situation. The input is the call instruction from the server, and the output is connecting the phone and explaining the situation to the other party.

[0454] Step 7:

[0455] The server uses a generative AI model to create a first aid guide based on the user's emotional state and sends it to the device. The device then provides this guide to the user. Even if communication is unstable, the device continues providing guidance using the installed language model. The input is the user's emotional state and situation data, and the output is the first aid guide content.

[0456] Specific examples

[0457] The following prompts can be fed into a generative AI model to generate specific emergency response instructions:

[0458] Example prompt:

[0459] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[0460] The above series of steps allows for a quick and appropriate emergency response while taking into account the user's emotional state.

[0461] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0462] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0463] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0464] [Second embodiment]

[0465] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0466] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0467] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0468] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0469] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0470] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0471] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0472] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0473] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0474] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0475] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0476] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0477] The present invention is a system for responding quickly and appropriately to emergencies. The system is built around a smartphone application that allows users who encounter an emergency to automatically contact the appropriate emergency services with the push of a button.

[0478] System configuration

[0479] This system consists of a user's smartphone device (hereinafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. A multimodal generation AI runs on the server side and analyzes the transmitted data.

[0480] System Operation

[0481] 1. Pressing the emergency button and collecting data

[0482] When a user presses the emergency button, the device activates the camera and starts recording video from that point onward, simultaneously recording audio data using the microphone, and obtaining the device's current location using the GPS module.

[0483] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[0484] 2. Data submission and analysis

[0485] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[0486] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[0487] 3. Determine emergency services and initiate a call

[0488] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[0489] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[0490] 4. First Aid and Navigation

[0491] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station, allowing the device to provide turn-by-turn navigation to the user.

[0492] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[0493] This system will enable users to respond quickly and accurately in emergency situations, which is expected to improve the chances of survival. In addition, by utilizing speech analysis methods and language models that work even in unstable communication environments, appropriate support can be provided in any situation.

[0494] The processing flow will be explained below.

[0495] Step 1:

[0496] When a user presses the emergency button, the device immediately activates the camera and starts recording video. At the same time, the microphone is activated to record surrounding audio. The device also uses the GPS module to obtain its current location.

[0497] Step 2:

[0498] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[0499] Step 3:

[0500] The server adds the received data packets to the analysis queue and activates the multimodal generation AI. The server analyzes the video data to identify emergency situations (e.g., traffic accidents, fires, collapsed people, etc.) and analyzes the surrounding sound environment from the audio data.

[0501] Step 4:

[0502] Based on the analysis results, the server determines the type and severity of the emergency, identifies the current geographic location from the location information, and lists the nearest emergency services (e.g., police station, fire station, hospital, etc.).

[0503] Step 5:

[0504] The server selects the most appropriate emergency service contact based on the determined emergency situation and returns that information to the device.

[0505] Step 6:

[0506] The device responds to the received instructions and automatically calls the selected emergency service contact, and plays a voice message generated by the server explaining the situation.

[0507] Step 7:

[0508] After the call is made, the device will begin providing first aid instructions to the user. The server will calculate the optimal route to the nearest hospital or fire station based on the user's location and send that information to the device.

[0509] Step 8:

[0510] The device will display the calculated optimal route information to the user, initiate turn-by-turn navigation, and even if the connection is unstable, it can continue to provide first aid instructions using pre-installed language models.

[0511] Step 9:

[0512] By following the first aid and navigation instructions, users can facilitate appropriate responses at the scene until emergency services arrive, which can lead to faster and more accurate responses and potentially improved survival rates.

[0513] Example 1

[0514] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0515] When encountering an emergency, it can be difficult for individuals to contact emergency services quickly and appropriately. This is especially true when users are in a state of panic or have difficulty explaining the situation, which can delay appropriate responses. Furthermore, there is a risk that damage may escalate due to users not knowing how to respond appropriately. There is a need for a system that can resolve these issues and provide a fast and appropriate emergency response.

[0516] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0517] In this invention, the server includes means for activating an image acquisition means and a voice acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, voice data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for calculating an optimal route based on the user's location information and providing navigation, thereby enabling the user to respond to an emergency quickly and accurately.

[0518] "Image acquisition means" refers to a device or function that uses a camera or other image sensor to acquire video or image information about the surroundings.

[0519] An "audio capture means" is a device or function that uses a microphone or other audio sensor to record or capture ambient sounds.

[0520] "Location Information Acquisition Means" means a device or function that acquires current geographic location information using GPS or other location determination technology.

[0521] The "transmission means" is a function that assembles the acquired data into a single packet and transmits it via the Internet or other communication means.

[0522] "Analysis means" refers to software or hardware that analyzes received data and identifies specific conditions or information.

[0523] "Multimodal generation AI" is an artificial intelligence technology that integrates and analyzes multiple modals (input formats) such as images, audio, and text.

[0524] "Emergency Contacts" are contact details for agencies and services (e.g., ambulance, police, fire department, etc.) that will best respond to a particular emergency.

[0525] "Contact methods" are devices or functions that automatically initiate a call to an emergency contact and convey the necessary information.

[0526] The "guidance means" is a function for providing first aid and navigation information to the user.

[0527] "Navigation means" is a function that calculates the optimal route based on the user's location information and provides real-time directions.

[0528] MODE FOR CARRYING OUT THE INVENTION

[0529] The present invention is a system for responding promptly and appropriately to an emergency situation, and functions in cooperation with a user, a terminal, and a server. A specific embodiment of this system will be described below.

[0530] 1. Components and Hardware Configuration

[0531] Terminal

[0532] The device is a smartphone owned by the user and is equipped with the following hardware:

[0533] Camera: Captures video and images of the surrounding area.

[0534] Microphone: Records surrounding sounds.

[0535] GPS module: Obtains current location information.

[0536] These hardware components function as an "image acquisition means," "audio acquisition means," and "location information acquisition means." In addition, a dedicated application is installed, and these functions are automatically activated when the emergency button is pressed.

[0537] server

[0538] The server will be deployed in a cloud environment and will utilize the following software and technologies:

[0539] Multimodal generative AI: Integrated analysis of video, audio, and location data to identify emergencies.

[0540] Communication means: Receives data sent from the terminal and sends the analysis results to the terminal.

[0541] 2. Usage and Data Processing

[0542] Pressing the emergency button

[0543] When a user presses the emergency button within a smartphone application, the device automatically activates the camera, microphone, and GPS to collect data.

[0544] Sending data

[0545] The terminal combines the collected video data, audio data, and location information into a single data packet, which is then transmitted to a server via the Internet by a transmitting means.

[0546] Data analysis

[0547] The server analyzes the received data packets using an analytical method (multimodal generative AI), which identifies the type of emergency and selects the appropriate emergency services (ambulance, police, fire department, etc.).

[0548] Example: If the server analyzes the data it receives and determines that it is a "traffic accident" based on video and audio, the analysis results will also include the extent of injuries and the danger level of the scene.

[0549] Starting a call

[0550] The device receives instructions from the server and automatically calls the designated emergency service, explaining the situation using a voice message generated by the server.

[0551] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, providing details of the emergency situation via voice.

[0552] First Aid Guide and Navigation

[0553] While waiting for the ambulance to arrive, the device provides a pre-prepared first aid guide from the server. The server also tracks the user's location in real time and calculates the optimal route to the nearest hospital or fire station. As a navigation tool, the device provides turn-by-turn navigation information to the user.

[0554] Example: The device provides voice and text instructions for CPR and displays directions to the nearest hospital.

[0555] Prompt Sentence Examples

[0556] An example of a prompt to be input to the generative AI model when a user encounters an emergency:

[0557] "I was involved in a traffic accident and someone was injured. My current location is latitude xx.xxxx, longitude yy.yyyy."

[0558] "A fire has broken out and there are many injured people nearby. Your current location is latitude xx.xxxx, longitude yy.yyyy."

[0559] In this way, the entire system works together to provide comprehensive support in emergencies, enabling users to quickly take the most appropriate action.

[0560] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0561] Step 1:

[0562] User operations

[0563] The user presses the emergency button in the application.

[0564] Input: User touch input.

[0565] Output: Emergency button pressed event.

[0566] Specific action: The user taps the emergency button displayed on the smartphone screen.

[0567] Step 2:

[0568] Device behavior: Data collection

[0569] The device detects when the emergency button is pressed and activates the camera, microphone, and GPS module to collect data.

[0570] Input: Emergency button press event.

[0571] Output: Video data, audio data, location information.

[0572] Specific behavior:

[0573] The camera will activate and record video of the emergency.

[0574] The microphone records the audio.

[0575] The GPS module obtains the current location information.

[0576] Step 3:

[0577] Device operation: Data transmission

[0578] The device bundles the collected data (video data, audio data, location information) into a single data packet and sends it to the server.

[0579] Input: Video data, audio data, location information.

[0580] Output: Data packets sent to the server.

[0581] Specific behavior:

[0582] Video, audio, and location information are packaged into packets.

[0583] The packet is sent over the Internet to a server.

[0584] Step 4:

[0585] Server Operation: Data Analysis

[0586] The server analyzes the received data packets and runs a multimodal generative AI to identify emergencies.

[0587] Input: Data packet.

[0588] Output: Emergency type, detailed analysis results.

[0589] Specific behavior:

[0590] The received video data, audio data, and location information are expanded.

[0591] Multimodal generative AI analyzes the data, identifies the type of emergency (e.g., traffic accident, fire, sudden illness, etc.), and determines the detailed situation.

[0592] Step 5:

[0593] Server Behavior: Emergency Services Decision

[0594] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, or fire department).

[0595] Input: Emergency type, detailed analysis results.

[0596] Output: Emergency services contact information.

[0597] Specific behavior:

[0598] Retrieve the best contact information from a database for the type of emergency.

[0599] Send contact information for selected emergency services to the device.

[0600] Step 6:

[0601] Device Action: Initiating a call

[0602] The device receives instructions from the server and automatically calls emergency services. During the call, it plays a voice message generated by the server explaining the situation.

[0603] Input: Emergency services contact information, server-generated voice message.

[0604] Output: Calls to emergency services, playback of voice messages.

[0605] Specific behavior:

[0606] Your device will automatically dial the emergency services number.

[0607] Once the call is connected, a server-generated voice message is played, detailing the emergency.

[0608] Step 7:

[0609] Device Operation: First Aid and Navigation

[0610] While waiting for the ambulance to arrive, the device will provide the user with first aid guidance and navigate the optimal route to the nearest hospital or fire station.

[0611] Input: First aid guide information provided by the server, real-time location information.

[0612] Output: First aid guide, navigation information.

[0613] Specific behavior:

[0614] The device provides the user with pre-programmed first aid instructions (audio, text, video).

[0615] The server calculates the optimal route based on the user's location information and sends navigation information to the terminal.

[0616] The device displays real-time turn-by-turn navigation to the user.

[0617] (Application example 1)

[0618] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0619] Conventional emergency response systems have the drawback of making it difficult for users to respond quickly and appropriately when they encounter an emergency. In particular, they require complex procedures, such as initiating a call to an emergency contact and providing first aid instructions, which often prevents users from responding calmly. Furthermore, they lack the functionality to provide an optimal route based on the user's current location, making it difficult to evacuate to a safe location efficiently. Therefore, there is a need for a system that supports users' quick and accurate actions in emergency response.

[0620] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0621] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for the guidance means to calculate an optimal travel route based on the user's current location and provide navigation, thereby enabling the user to respond quickly and accurately to an emergency situation and efficiently evacuate to a safe place.

[0622] - A "trigger" is an input device such as a button or switch that a user operates to initiate a specific action.

[0623] "Image acquisition means" refers to a device that collects video data using a camera or the like.

[0624] "Location information acquisition means" is a device that identifies the current location using a GPS module or the like.

[0625] "Audio data" refers to a recording of sound captured using a microphone or the like.

[0626] "Transmission means" means a device or process that transmits collected data to a server via the Internet or other communication means.

[0627] "Analysis means" refers to algorithms or software that process received data and identify emergency conditions.

[0628] A "call" is a means of making a voice communication, which may be made over a telephone line or the Internet.

[0629] "Contact Assistance" means any device or software that automatically places a call to the appropriate emergency service.

[0630] A "guiding means" is a device or software that provides useful information to a user.

[0631] A "First Aid Guide" is information that provides instructions on how to provide basic medical assistance in an emergency.

[0632] "Navigation" is a route guidance function for guiding a user to a specific destination.

[0633] "Current location of user" is real-time location information of the user identified by the location information acquisition means.

[0634] This invention is a system for responding quickly and accurately to emergencies, and is built around a user's smartphone terminal. The configuration and operation of the system are described in detail below.

[0635] 1. System Configuration

[0636] This system consists of a user's smartphone terminal (hereinafter referred to as the terminal), a server connected via the Internet, and an emergency service. The terminal is equipped with a camera, microphone, GPS, and the application of this invention. A multimodal generation AI runs on the server side.

[0637] 2. System Operation

[0638] Pressing the emergency button and collecting data

[0639] When the user presses the emergency button, the application installed on the device activates the camera and starts recording video from that point. At the same time, it also records audio data using the microphone. It also obtains the current location information using the GPS module.

[0640] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[0641] Data transmission and analysis

[0642] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[0643] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[0644] Determining emergency services and initiating a call

[0645] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[0646] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[0647] First Aid and Navigation

[0648] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station based on the user's current location, allowing the device to provide turn-by-turn navigation to the user.

[0649] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[0650] Prompt Sentence Examples

[0651] This is an emergency. A user has suddenly lost consciousness at home. Location, video, and audio data are attached below. Please take appropriate action.

[0652] Location: Tokyo, Japan

[0653] Video data: emergency_video.avi

[0654] Audio data: emergency_audio.wav

[0655] Hardware and software used

[0656] Device: Smartphone (camera, microphone, GPS)

[0657] Server: Internet Server

[0658] Software: Multimodal generative AI, Python, OpenCV, PyAudio, Geopy, Requests library

[0659] This allows users to respond quickly and accurately to emergencies and efficiently evacuate to a safe location. In addition, the server tracks users' locations in real time and provides first aid guidance and optimal route navigation, which is expected to improve survival rates.

[0660] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0661] Step 1:

[0662] The device detects the pressing of the emergency button. The input is the user pressing the button, and the output is the activation of the camera, microphone, and GPS. When the user performs a trigger operation, the device detects this event and activates the camera, microphone, and GPS sensors.

[0663] Step 2:

[0664] The device records video with the camera and audio with the microphone. The input is data information from the camera and microphone, and the output is a video file and an audio file. The device starts the camera, records video for a certain period of time (e.g., 10 seconds), and simultaneously records audio data using the microphone. These data are then saved to a file.

[0665] Step 3:

[0666] The device uses GPS to obtain current location information. The input is the signal from the GPS module and the output is location data. The device uses the GPS module to obtain real-time location information (latitude and longitude) and saves the data in the application.

[0667] Step 4:

[0668] The device collects video data, audio data, and location information and combines them into packets. The input is each file and location data, and the output is a data packet. The device combines the video file, audio file, and location information into a single packet.

[0669] Step 5:

[0670] The terminal sends a data packet to the server. The input is the data packet, and the output is a notification of successful data transmission to the server. The terminal sends the data packet to the server via the Internet and confirms its success.

[0671] Step 6:

[0672] The server analyzes the received data. The input is a data packet, and the output is the emergency identification result. The server separates video, audio, and location information from the received data packet, analyzes this data using a generative AI model, and identifies the type of emergency.

[0673] Step 7:

[0674] The server determines the most appropriate emergency contact based on the analysis results. The input is the analysis results, and the output is emergency contact information. Based on the analysis results, the server selects the most appropriate emergency service (ambulance, police, fire department, etc.).

[0675] Step 8:

[0676] The device receives instructions from the server and automatically calls the appropriate emergency service. The input is emergency contact information, and the output is a successful call. The device automatically calls the emergency contact provided by the server and starts the call.

[0677] Step 9:

[0678] The device automatically explains the situation using a voice message generated by the server. The input is the voice message from the server, and the output is the completion of sending the voice message. The device plays the voice message generated by the server and explains the situation to the emergency services.

[0679] Step 10:

[0680] The terminal provides the user with a first aid guide prepared in advance on the server side. The input is first aid guide information, and the output is completion of guide provision. The terminal provides the first aid guide provided by the server to the user in voice or text format.

[0681] Step 11:

[0682] The server tracks the user's real-time location and calculates the optimal route to the nearest hospital or fire station. The input is the user's current location information and the output is the optimal route information. The server uses GPS data to identify the user's current location and calculates the optimal route to the nearest emergency facility.

[0683] Step 12:

[0684] The terminal provides optimal route navigation to the user. The input is optimal route information, and the output is navigation completion. The terminal provides turn-by-turn navigation to the user based on the optimal route information provided by the server.

[0685] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0686] The present invention is a system that takes into consideration the user's emotions in an emergency and can respond quickly and appropriately. This system provides a means for responding to an emergency through a smartphone app, and provides appropriate guidance according to the user's stress and impatience.

[0687] System configuration

[0688] This system consists of the user's smartphone device (hereafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. The server runs a multimodal generation AI and an emotion recognition engine, which analyzes the transmitted data.

[0689] System Operation

[0690] 1. Pressing the emergency button and collecting data

[0691] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously enables the microphone to record audio, acquires location information using the GPS module, and packetizes the audio data for transmission to the server.

[0692] Example: When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and obtaining location information.

[0693] 2. Data submission and analysis

[0694] The device then combines the captured video and audio files, along with location information, into a single data packet and sends it to the server, which then queues the data for analysis and activates the multimodal generation AI and emotion recognition engine.

[0695] Example: The server analyzes the data it receives, identifies traffic accidents and emergencies from voices, and recognizes emotions from the user's tone of voice.

[0696] 3. Emotion Recognition and Emergency Service Decisions

[0697] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as anxiety, fear, panic, etc. The server then uses the analysis results to determine the type and severity of the emergency and selects the appropriate emergency services (e.g., police, ambulance, fire department, etc.).

[0698] Example: If the server detects that the user is panicking, it will determine the level of urgency as high and send instructions to the device to immediately contact an ambulance and the police.

[0699] 4. Automatic contact and first aid provision

[0700] The device receives instructions from the server and automatically calls the selected emergency service contact. The device plays a voice message generated by the server, explaining the situation. Furthermore, the device adjusts the first aid guidance content according to the user's emotional state and provides it to the user.

[0701] Example: For a user in a panicked state, a calming voice will provide instructions on CPR and how to treat injuries.

[0702] 5. Providing location information and navigation

[0703] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information, and the device displays this information to the user and provides turn-by-turn navigation.

[0704] Examples include providing detailed directions to the nearest hospital to help users avoid getting lost, and continuing to provide first aid instructions even when communication is unstable, using pre-installed language models.

[0705] 6. Providing Additional Information

[0706] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions, so that emergency services can make appropriate preparations before arriving on scene.

[0707] Example: If the user is very disoriented, this information can also be communicated to emergency responders so that they can provide psychological support when they arrive on scene.

[0708] In this way, the system, including the emotion recognition engine, enables a fast and accurate response in emergency situations, which is expected to increase the user's sense of security and improve the overall survival rate.

[0709] The processing flow will be explained below.

[0710] Step 1:

[0711] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously activates the microphone to capture surrounding audio, and uses the GPS module to obtain the device's current location.

[0712] Step 2:

[0713] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[0714] Step 3:

[0715] The server adds the received data packets to an analysis queue and activates a multimodal generation AI and emotion recognition engine, which analyzes video data to identify emergency situations and recognizes user emotions from audio data.

[0716] Step 4:

[0717] The server's emotion recognition engine analyzes the user's tone of voice, speed, and word choice from the audio data to determine whether the user is anxious, calm, or panicked.

[0718] Step 5:

[0719] Based on the analysis results, the server determines the type and severity of the emergency, selects the most appropriate emergency service contact, and sends this information back to the device.

[0720] Step 6:

[0721] Based on the instructions received, the device automatically calls the selected emergency service contact and plays a server-generated voice message explaining the situation.

[0722] Step 7:

[0723] The device provides first aid instructions to the user. If the server's emotion recognition engine determines that the user's emotional state is impatient or panicked, the device adjusts the content and speed of the instructions to help the user calm down.

[0724] Step 8:

[0725] The server calculates the optimal route to the nearest hospital or fire station based on real-time location information and sends that information to the device, which then displays the calculated route information to the user and begins turn-by-turn navigation.

[0726] Step 9:

[0727] The emotion recognition engine takes into account the user's emotions and provides additional information to emergency contacts, allowing them to make appropriate preparations before emergency services arrive on scene.

[0728] Step 10:

[0729] By following the instructions on the device and receiving first aid and navigation, users can take appropriate action until emergency services arrive, which will enable quick and accurate response and is expected to improve the survival rate.

[0730] Example 2

[0731] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0732] Conventional emergency response systems are required to respond not only based on user reports but also by taking into account the user's emotional state. However, current systems have difficulty accurately recognizing the user's emotions and providing appropriate guidance and emergency measures. Another challenge is providing accurate responses in environments with unstable communication conditions.

[0733] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0734] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for adjusting first aid guidance according to the user's emotional state, and a means for calculating an optimal route based on the user's real-time location information and providing navigation, thereby enabling a quick and accurate emergency response that also takes the user's emotional state into consideration.

[0735] 1. "Image capture means" refers to a device that activates a camera and captures video or still images when a user presses a trigger.

[0736] 2. "Location information acquisition means" refers to a device that acquires location information of the user's current location using a GPS module or the like.

[0737] 3. "Transmitting means" refers to a device that packetizes the acquired image data, audio data, and location information and transmits them to the server.

[0738] 4. "Analysis means" is a device that analyzes the data transmitted by the transmission means and identifies an emergency situation or the user's emotions.

[0739] 5. "Contact Means" means a device that determines the appropriate emergency contact based on the emergency situation identified by the Analysis Means and automatically initiates a call.

[0740] 6. "Guidance means" is a device that provides the user with first aid guidance and appropriate actions according to their emotional state.

[0741] 7. "Means for providing navigation" means a device that calculates the optimal route based on the user's real-time location information and provides turn-by-turn navigation.

[0742] 8. "Acoustic analysis means" refers to a device that analyzes the surrounding acoustic environment and the user's emotions based on audio data.

[0743] 9. "Language model" refers to a model that is installed to function even in environments with unstable communication and is used to generate emergency and guidance messages.

[0744] MODE FOR CARRYING OUT THE INVENTION

[0745] This invention is an emergency response system that takes into account the user's emotions and can respond quickly and appropriately in an emergency. This system operates via a smartphone app and handles emergencies using an emotion recognition engine and multimodal generation AI.

[0746] Hardware Configuration

[0747] The hardware configuration of this system is as follows:

[0748] 1. Device: A smartphone carried by the user, equipped with the following functions:

[0749] Camera features

[0750] Microphone function

[0751] GPS Modules

[0752] Internet communication function

[0753] 2. Server: A back-end system connected to terminals via the Internet, and includes the following components:

[0754] Multimodal generation AI (audio data analysis, image data analysis)

[0755] Emotion Recognition Engine

[0756] Data Analysis Engine

[0757] Software Configuration

[0758] 1. Terminal application: A smartphone application used by the user that performs the following actions in the event of an emergency.

[0759] Emergency button activates camera and microphone

[0760] Obtaining location information

[0761] Packetizing acquired data (video, audio, location information) and sending it to the server

[0762] 2. Server application: Software that runs on the server side and provides the following functions:

[0763] Analyzing received data

[0764] Recognizing the user's emotional state

[0765] Determining the type and severity of an emergency

[0766] Selection of emergency contacts and automatic call instructions

[0767] Generate and submit a first aid guide

[0768] Real-time location-based navigation

[0769] Example of operation

[0770] To give a concrete example of how this works, consider the following scenario:

[0771] 1. If the emergency button is pressed:

[0772] A user encounters a traffic accident and presses the emergency button on their smartphone.

[0773] The device immediately activates the camera and records video of the accident scene.

[0774] At the same time, the microphone is enabled to record audio and location information is obtained using GPS.

[0775] The terminal transmits this data to the server.

[0776] 2. Data transmission and analysis:

[0777] The server adds the received video, audio, and location information to an analysis queue and activates the multimodal generation AI and emotion recognition engine.

[0778] The server analyzes the user's emotional state (e.g., panic) from the tone of the voice and determines the appropriate emergency response.

[0779] 3. Contacting the appropriate emergency services:

[0780] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and issues an automatic call instruction.

[0781] The device automatically contacts emergency services and plays a server-generated voice message explaining the situation.

[0782] Prompt Sentence Examples

[0783] Examples of prompts include:

[0784] "A user encounters a traffic accident and presses the emergency button. The mobile device collects video, audio, and GPS data and sends it to a server. The server analyzes the urgency of the accident and coordinates the appropriate emergency services. Please explain how the system works."

[0785] This system enables a fast and effective response in emergency situations while taking into consideration the user's feelings, which is expected to improve the user's sense of security and increase the overall survival rate.

[0786] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0787] Step 1:

[0788] The user presses the emergency button.

[0789] Input: The user presses the emergency button.

[0790] Output: The emergency button press signal is transmitted to the terminal.

[0791] The device immediately activates the camera and starts recording video, while simultaneously activating the microphone for audio recording and the GPS module for location information.

[0792] Specific operation: Video recording, audio recording, and location information acquisition of the accident scene are carried out simultaneously.

[0793] Step 2:

[0794] The terminal combines the acquired video file, audio file, and location information into a single data packet and sends it to the server.

[0795] Input: Video data acquired by the camera, audio data recorded by the microphone, and location information acquired by GPS.

[0796] Output: A data packet is generated and sent to the server.

[0797] Specific operation: Various data is packetized and sent to a server via the Internet.

[0798] Step 3:

[0799] The server adds the received data to a queue for analysis and processing, and activates the multimodal generation AI and emotion recognition engine.

[0800] Input: Data packets sent from the terminal.

[0801] Output: The analysis results include the type of emergency and the user's emotional state.

[0802] Specific operation: The received data is added to the analysis queue, and the AI ​​and emotion recognition engine are activated to begin data analysis.

[0803] Step 4:

[0804] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as impatience, fear, or panic.

[0805] Input: Audio data.

[0806] Output: Analysis results showing the user's emotional state.

[0807] Specific operation: Emotion analysis is performed using voice data to determine the user's emotional state.

[0808] Step 5:

[0809] Based on the analysis results, the server determines the type and urgency of the emergency and selects the appropriate emergency service (e.g., police, ambulance, fire department, etc.).

[0810] Input: The type of emergency situation as the analysis result and the user's emotional state.

[0811] Output: Instructions for optimal emergency service selection.

[0812] Specific operation: Based on the analysis results, the level of urgency is determined to be high, and appropriate emergency services are immediately selected.

[0813] Step 6:

[0814] The device receives instructions from the server and automatically calls the selected emergency service contact.

[0815] Input: Instructions from the server.

[0816] Output: Initiation of automatic contact with emergency services.

[0817] Specific operation: The device makes an automatic call and plays a voice message generated by the server to explain the situation.

[0818] Step 7:

[0819] The device provides first aid guidance according to the user's emotional state.

[0820] Input: Analysis results indicating the user's emotional state.

[0821] Output: First aid guide depending on emotional state.

[0822] Specific actions: Provide breathing techniques and first aid instructions in a calm voice.

[0823] Step 8:

[0824] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information.

[0825] Input: The user's real-time location.

[0826] Output: Optimal route information.

[0827] Specific operation: Analyze real-time location information and calculate the optimal route.

[0828] Step 9:

[0829] The device displays optimal route information to the user and provides turn-by-turn navigation.

[0830] Input: Optimal route information.

[0831] Output: Navigation guide.

[0832] Specific behavior: Support the user by providing detailed directions to the nearest hospital.

[0833] Step 10:

[0834] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions.

[0835] Input: Data indicating the user's emotional state.

[0836] Output: Additional information for emergency contacts.

[0837] Specific behavior: If the user is confused, this information will also be conveyed to emergency responders so that psychological support can be prepared when they arrive at the scene.

[0838] (Application example 2)

[0839] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0840] In the event of an emergency, if a user feels anxious or scared, it can be difficult to respond appropriately. Furthermore, conventional emergency response systems do not take the user's emotional state into account, which can lead to inadequate responses. Furthermore, the system must function even in environments with unstable communications. The present invention aims to solve these problems and provide a system that allows users to respond to emergencies with peace of mind.

[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0842] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, an emotion recognition means for analyzing the user's emotional state identified by the analysis means, a contact means for determining an appropriate emergency contact based on the analysis result of the emotion recognition means and automatically initiating a call, and a guidance means for providing guidance on appropriate first aid according to the user's emotional state. This enables a quick and appropriate emergency response while taking the user's emotional state into consideration.

[0843] "Image capture means" refers to a technical means for capturing an image or video using a device such as a camera when a user activates a trigger.

[0844] "Location information acquisition means" refers to a technical means for determining the current location of a terminal using GPS or other positioning systems.

[0845] The "transmission means" is a means having a function of transmitting the acquired image data, audio data, and location information to a server or other external system.

[0846] "Analysis Means" means the technical means used to identify an emergency situation and determine appropriate emergency contacts based on the Transmitted Data.

[0847] "Emotion recognition means" is a technical means for analyzing the user's emotional state from voice data, etc., and identifying emotions such as impatience or fear.

[0848] A "contact method" is a method that has the ability to automatically initiate a call to an emergency contact based on an identified emergency condition and emotional state.

[0849] The "guidance means" is a technical means for providing a first aid guide in response to an emergency situation faced by a user, taking into consideration the user's emotional state.

[0850] MODE FOR CARRYING OUT THE INVENTION

[0851] To implement this invention, the user's smartphone terminal, the server, and the emergency service work together. The hardware and software used in each step and the processing content thereof will be specifically described below.

[0852] System configuration

[0853] 1. Hardware

[0854] Device (smartphone): Smartphones are equipped with a camera, microphone, and GPS module. These functions are used to acquire images, audio, and location information.

[0855] Server: The server runs a multimodal generation AI and emotion recognition engine, which analyzes the transmitted data.

[0856] 2. Software

[0857] Image capture tool: Software that activates the camera and captures images and videos.

[0858] Audio recording software: Records audio using a microphone and generates audio data.

[0859] Location information acquisition software: Uses the GPS module to determine the device's current location.

[0860] Data transmission software: Collects the acquired data into a single data packet and sends it to a server over the Internet.

[0861] Emotion recognition engine: Analyzes the user's emotional state from voice data and identifies emotions such as impatience and fear. Specifically, it uses an emotion recognition model from the Transformers library.

[0862] Multimodal generative AI: Identifies emergencies based on submitted data and generates appropriate responses. Uses the nlptown / bert-base-multilingual-uncased-sentiment model.

[0863] Contact Method: Automatically call emergency contacts based on the analyzed results.

[0864] Guidance: Provides first aid guidance according to the user's emotional state. Even if communication is unstable, the guidance continues using a pre-installed language model.

[0865] System Operation

[0866] 1. Pressing the emergency button and collecting data

[0867] When a user presses the emergency button, the device's camera automatically activates and begins recording images and videos. At the same time, the microphone activates and records audio. The GPS module also activates and collects location information. This data is packetized within the device and sent to the server.

[0868] 2. Data Analysis and Emotion Recognition

[0869] The server analyzes the received data and identifies the user's emotional state from the user's voice tone and environmental sounds. An emotion recognition engine analyzes the voice data and identifies whether the user is expressing emotions such as anxiety or fear, and determines the appropriate emergency measures.

[0870] Examples:

[0871] When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and acquiring location information. The server analyzes this data to understand the circumstances of the accident and the user's emotional state.

[0872] 3. Selecting and contacting emergency services

[0873] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and automatically calls the contacts. The server generates a voice message explaining the situation.

[0874] 4. First aid information

[0875] The system takes into account the user's emotional state and provides the most appropriate first aid guidance. For example, if a user is in a panic, a slow voice will guide them through CPR and how to treat injuries.

[0876] Examples:

[0877] By inputting the following prompts into the generative AI model, specific emergency response instructions can be generated.

[0878] Example prompt:

[0879] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[0880] As a result, it is possible to take prompt and appropriate emergency measures while taking into consideration the emotional state of the user.

[0881] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0882] Step 1:

[0883] When a user presses the emergency button, the device immediately activates the camera and starts recording images and videos. At the same time, it activates the microphone to record audio and uses the GPS module to obtain location information. These data are packetized and ready to be transmitted. The input is the user pressing the button, and the output is image data, audio data, and location information.

[0884] Step 2:

[0885] The device collects the acquired image data, audio data, and location information into a single data packet and sends it to a server via the Internet. The input is the image data, audio data, and location information, and the output is the data packet sent to the server.

[0886] Step 3:

[0887] The server analyzes the received data packets using a multimodal generation AI and emotion recognition engine. First, it identifies the emergency situation from image data and location information, and then analyzes the audio data to recognize the user's emotional state. The input is the data packets, and the output is the emergency situation identification result and the user's emotional state recognition result.

[0888] Step 4:

[0889] The server analyzes the voice data using an emotion recognition engine to determine whether the user is expressing emotions such as anxiety or fear. For example, it uses an emotion recognition model from the transformers library. The input is the voice data, and the output is the identification of the emotional state.

[0890] Step 5:

[0891] The server determines the appropriate emergency contact based on the analysis results and instructs the device to automatically initiate a call. The server generates the necessary call content using a generative AI model and sends it to the device. The input is the emergency state and emotional state identification results, and the output is the emergency contact and call content.

[0892] Step 6:

[0893] The device receives instructions from the server and automatically calls the appropriate emergency contact. It responds by playing a voice message generated by the server and explaining the situation. The input is the call instruction from the server, and the output is connecting the phone and explaining the situation to the other party.

[0894] Step 7:

[0895] The server uses a generative AI model to create a first aid guide based on the user's emotional state and sends it to the device. The device then provides this guide to the user. Even if communication is unstable, the device continues providing guidance using the installed language model. The input is the user's emotional state and situation data, and the output is the first aid guide content.

[0896] Specific examples

[0897] The following prompts can be fed into a generative AI model to generate specific emergency response instructions:

[0898] Example prompt:

[0899] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[0900] The above series of steps allows for a quick and appropriate emergency response while taking into account the user's emotional state.

[0901] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0902] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0903] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0904] [Third embodiment]

[0905] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0906] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0907] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0908] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0909] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0910] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0911] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0912] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0913] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0914] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0915] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0916] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0917] The present invention is a system for responding quickly and appropriately to emergencies. The system is built around a smartphone application that allows users who encounter an emergency to automatically contact the appropriate emergency services with the push of a button.

[0918] System configuration

[0919] This system consists of a user's smartphone device (hereinafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. A multimodal generation AI runs on the server side and analyzes the transmitted data.

[0920] System Operation

[0921] 1. Pressing the emergency button and collecting data

[0922] When a user presses the emergency button, the device activates the camera and starts recording video from that point onward, simultaneously recording audio data using the microphone, and obtaining the device's current location using the GPS module.

[0923] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[0924] 2. Data submission and analysis

[0925] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[0926] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[0927] 3. Determine emergency services and initiate a call

[0928] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[0929] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[0930] 4. First Aid and Navigation

[0931] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station, allowing the device to provide turn-by-turn navigation to the user.

[0932] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[0933] This system will enable users to respond quickly and accurately in emergency situations, which is expected to improve the chances of survival. In addition, by utilizing speech analysis methods and language models that work even in unstable communication environments, appropriate support can be provided in any situation.

[0934] The processing flow will be explained below.

[0935] Step 1:

[0936] When a user presses the emergency button, the device immediately activates the camera and starts recording video. At the same time, the microphone is activated to record surrounding audio. The device also uses the GPS module to obtain its current location.

[0937] Step 2:

[0938] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[0939] Step 3:

[0940] The server adds the received data packets to the analysis queue and activates the multimodal generation AI. The server analyzes the video data to identify emergency situations (e.g., traffic accidents, fires, collapsed people, etc.) and analyzes the surrounding sound environment from the audio data.

[0941] Step 4:

[0942] Based on the analysis results, the server determines the type and severity of the emergency, identifies the current geographic location from the location information, and lists the nearest emergency services (e.g., police station, fire station, hospital, etc.).

[0943] Step 5:

[0944] The server selects the most appropriate emergency service contact based on the determined emergency situation and returns that information to the device.

[0945] Step 6:

[0946] The device responds to the received instructions and automatically calls the selected emergency service contact, and plays a voice message generated by the server explaining the situation.

[0947] Step 7:

[0948] After the call is made, the device will begin providing first aid instructions to the user. The server will calculate the optimal route to the nearest hospital or fire station based on the user's location and send that information to the device.

[0949] Step 8:

[0950] The device will display the calculated optimal route information to the user, initiate turn-by-turn navigation, and even if the connection is unstable, it can continue to provide first aid instructions using pre-installed language models.

[0951] Step 9:

[0952] By following the first aid and navigation instructions, users can facilitate appropriate responses at the scene until emergency services arrive, which can lead to faster and more accurate responses and potentially improved survival rates.

[0953] Example 1

[0954] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0955] When encountering an emergency, it can be difficult for individuals to contact emergency services quickly and appropriately. This is especially true when users are in a state of panic or have difficulty explaining the situation, which can delay appropriate responses. Furthermore, there is a risk that damage may escalate due to users not knowing how to respond appropriately. There is a need for a system that can resolve these issues and provide a fast and appropriate emergency response.

[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0957] In this invention, the server includes means for activating an image acquisition means and a voice acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, voice data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for calculating an optimal route based on the user's location information and providing navigation, thereby enabling the user to respond to an emergency quickly and accurately.

[0958] "Image acquisition means" refers to a device or function that uses a camera or other image sensor to acquire video or image information about the surroundings.

[0959] An "audio capture means" is a device or function that uses a microphone or other audio sensor to record or capture ambient sounds.

[0960] "Location Information Acquisition Means" means a device or function that acquires current geographic location information using GPS or other location determination technology.

[0961] The "transmission means" is a function that assembles the acquired data into a single packet and transmits it via the Internet or other communication means.

[0962] "Analysis means" refers to software or hardware that analyzes received data and identifies specific conditions or information.

[0963] "Multimodal generation AI" is an artificial intelligence technology that integrates and analyzes multiple modals (input formats) such as images, audio, and text.

[0964] "Emergency Contacts" are contact details for agencies and services (e.g., ambulance, police, fire department, etc.) that will best respond to a particular emergency.

[0965] "Contact methods" are devices or functions that automatically initiate a call to an emergency contact and convey the necessary information.

[0966] The "guidance means" is a function for providing first aid and navigation information to the user.

[0967] "Navigation means" is a function that calculates the optimal route based on the user's location information and provides real-time directions.

[0968] MODE FOR CARRYING OUT THE INVENTION

[0969] The present invention is a system for responding promptly and appropriately to an emergency situation, and functions in cooperation with a user, a terminal, and a server. A specific embodiment of this system will be described below.

[0970] 1. Components and Hardware Configuration

[0971] Terminal

[0972] The device is a smartphone owned by the user and is equipped with the following hardware:

[0973] Camera: Captures video and images of the surrounding area.

[0974] Microphone: Records surrounding sounds.

[0975] GPS module: Obtains current location information.

[0976] These hardware components function as an "image acquisition means," "audio acquisition means," and "location information acquisition means." In addition, a dedicated application is installed, and these functions are automatically activated when the emergency button is pressed.

[0977] server

[0978] The server will be deployed in a cloud environment and will utilize the following software and technologies:

[0979] Multimodal generative AI: Integrated analysis of video, audio, and location data to identify emergencies.

[0980] Communication means: Receives data sent from the terminal and sends the analysis results to the terminal.

[0981] 2. Usage and Data Processing

[0982] Pressing the emergency button

[0983] When a user presses the emergency button within a smartphone application, the device automatically activates the camera, microphone, and GPS to collect data.

[0984] Sending data

[0985] The terminal combines the collected video data, audio data, and location information into a single data packet, which is then transmitted to a server via the Internet by a transmitting means.

[0986] Data analysis

[0987] The server analyzes the received data packets using an analytical method (multimodal generative AI), which identifies the type of emergency and selects the appropriate emergency services (ambulance, police, fire department, etc.).

[0988] Example: If the server analyzes the data it receives and determines that it is a "traffic accident" based on video and audio, the analysis results will also include the extent of injuries and the danger level of the scene.

[0989] Starting a call

[0990] The device receives instructions from the server and automatically calls the designated emergency service, explaining the situation using a voice message generated by the server.

[0991] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, providing details of the emergency situation via voice.

[0992] First Aid Guide and Navigation

[0993] While waiting for the ambulance to arrive, the device provides a pre-prepared first aid guide from the server. The server also tracks the user's location in real time and calculates the optimal route to the nearest hospital or fire station. As a navigation tool, the device provides turn-by-turn navigation information to the user.

[0994] Example: The device provides voice and text instructions for CPR and displays directions to the nearest hospital.

[0995] Prompt Sentence Examples

[0996] An example of a prompt to be input to the generative AI model when a user encounters an emergency:

[0997] "I was involved in a traffic accident and someone was injured. My current location is latitude xx.xxxx, longitude yy.yyyy."

[0998] "A fire has broken out and there are many injured people nearby. Your current location is latitude xx.xxxx, longitude yy.yyyy."

[0999] In this way, the entire system works together to provide comprehensive support in emergencies, enabling users to quickly take the most appropriate action.

[1000] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1001] Step 1:

[1002] User operations

[1003] The user presses the emergency button in the application.

[1004] Input: User touch input.

[1005] Output: Emergency button pressed event.

[1006] Specific action: The user taps the emergency button displayed on the smartphone screen.

[1007] Step 2:

[1008] Device behavior: Data collection

[1009] The device detects when the emergency button is pressed and activates the camera, microphone, and GPS module to collect data.

[1010] Input: Emergency button press event.

[1011] Output: Video data, audio data, location information.

[1012] Specific behavior:

[1013] The camera will activate and record video of the emergency.

[1014] The microphone records the audio.

[1015] The GPS module obtains the current location information.

[1016] Step 3:

[1017] Device operation: Data transmission

[1018] The device bundles the collected data (video data, audio data, location information) into a single data packet and sends it to the server.

[1019] Input: Video data, audio data, location information.

[1020] Output: Data packets sent to the server.

[1021] Specific behavior:

[1022] Video, audio, and location information are packaged into packets.

[1023] The packet is sent over the Internet to a server.

[1024] Step 4:

[1025] Server Operation: Data Analysis

[1026] The server analyzes the received data packets and runs a multimodal generative AI to identify emergencies.

[1027] Input: Data packet.

[1028] Output: Emergency type, detailed analysis results.

[1029] Specific behavior:

[1030] The received video data, audio data, and location information are expanded.

[1031] Multimodal generative AI analyzes the data, identifies the type of emergency (e.g., traffic accident, fire, sudden illness, etc.), and determines the detailed situation.

[1032] Step 5:

[1033] Server Behavior: Emergency Services Decision

[1034] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, or fire department).

[1035] Input: Emergency type, detailed analysis results.

[1036] Output: Emergency services contact information.

[1037] Specific behavior:

[1038] Retrieve the best contact information from a database for the type of emergency.

[1039] Send contact information for selected emergency services to the device.

[1040] Step 6:

[1041] Device Action: Initiating a call

[1042] The device receives instructions from the server and automatically calls emergency services. During the call, it plays a voice message generated by the server explaining the situation.

[1043] Input: Emergency services contact information, server-generated voice message.

[1044] Output: Calls to emergency services, playback of voice messages.

[1045] Specific behavior:

[1046] Your device will automatically dial the emergency services number.

[1047] Once the call is connected, a server-generated voice message is played, detailing the emergency.

[1048] Step 7:

[1049] Device Operation: First Aid and Navigation

[1050] While waiting for the ambulance to arrive, the device will provide the user with first aid guidance and navigate the optimal route to the nearest hospital or fire station.

[1051] Input: First aid guide information provided by the server, real-time location information.

[1052] Output: First aid guide, navigation information.

[1053] Specific behavior:

[1054] The device provides the user with pre-programmed first aid instructions (audio, text, video).

[1055] The server calculates the optimal route based on the user's location information and sends navigation information to the terminal.

[1056] The device displays real-time turn-by-turn navigation to the user.

[1057] (Application example 1)

[1058] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1059] Conventional emergency response systems have the drawback of making it difficult for users to respond quickly and appropriately when they encounter an emergency. In particular, they require complex procedures, such as initiating a call to an emergency contact and providing first aid instructions, which often prevents users from responding calmly. Furthermore, they lack the functionality to provide an optimal route based on the user's current location, making it difficult to evacuate to a safe location efficiently. Therefore, there is a need for a system that supports users' quick and accurate actions in emergency response.

[1060] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1061] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for the guidance means to calculate an optimal travel route based on the user's current location and provide navigation, thereby enabling the user to respond quickly and accurately to an emergency situation and efficiently evacuate to a safe place.

[1062] - A "trigger" is an input device such as a button or switch that a user operates to initiate a specific action.

[1063] "Image acquisition means" refers to a device that collects video data using a camera or the like.

[1064] "Location information acquisition means" is a device that identifies the current location using a GPS module or the like.

[1065] "Audio data" refers to a recording of sound captured using a microphone or the like.

[1066] "Transmission means" means a device or process that transmits collected data to a server via the Internet or other communication means.

[1067] "Analysis means" refers to algorithms or software that process received data and identify emergency conditions.

[1068] A "call" is a means of making a voice communication, which may be made over a telephone line or the Internet.

[1069] "Contact Assistance" means any device or software that automatically places a call to the appropriate emergency service.

[1070] A "guiding means" is a device or software that provides useful information to a user.

[1071] A "First Aid Guide" is information that provides instructions on how to provide basic medical assistance in an emergency.

[1072] "Navigation" is a route guidance function for guiding a user to a specific destination.

[1073] "Current location of user" is real-time location information of the user identified by the location information acquisition means.

[1074] This invention is a system for responding quickly and accurately to emergencies, and is built around a user's smartphone terminal. The configuration and operation of the system are described in detail below.

[1075] 1. System Configuration

[1076] This system consists of a user's smartphone terminal (hereinafter referred to as the terminal), a server connected via the Internet, and an emergency service. The terminal is equipped with a camera, microphone, GPS, and the application of this invention. A multimodal generation AI runs on the server side.

[1077] 2. System Operation

[1078] Pressing the emergency button and collecting data

[1079] When the user presses the emergency button, the application installed on the device activates the camera and starts recording video from that point. At the same time, it also records audio data using the microphone. It also obtains the current location information using the GPS module.

[1080] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[1081] Data transmission and analysis

[1082] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[1083] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[1084] Determining emergency services and initiating a call

[1085] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[1086] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[1087] First Aid and Navigation

[1088] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station based on the user's current location, allowing the device to provide turn-by-turn navigation to the user.

[1089] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[1090] Prompt Sentence Examples

[1091] This is an emergency. A user has suddenly lost consciousness at home. Location, video, and audio data are attached below. Please take appropriate action.

[1092] Location: Tokyo, Japan

[1093] Video data: emergency_video.avi

[1094] Audio data: emergency_audio.wav

[1095] Hardware and software used

[1096] Device: Smartphone (camera, microphone, GPS)

[1097] Server: Internet Server

[1098] Software: Multimodal generative AI, Python, OpenCV, PyAudio, Geopy, Requests library

[1099] This allows users to respond quickly and accurately to emergencies and efficiently evacuate to a safe location. In addition, the server tracks users' locations in real time and provides first aid guidance and optimal route navigation, which is expected to improve survival rates.

[1100] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1101] Step 1:

[1102] The device detects the pressing of the emergency button. The input is the user pressing the button, and the output is the activation of the camera, microphone, and GPS. When the user performs a trigger operation, the device detects this event and activates the camera, microphone, and GPS sensors.

[1103] Step 2:

[1104] The device records video with the camera and audio with the microphone. The input is data information from the camera and microphone, and the output is a video file and an audio file. The device starts the camera, records video for a certain period of time (e.g., 10 seconds), and simultaneously records audio data using the microphone. These data are then saved to a file.

[1105] Step 3:

[1106] The device uses GPS to obtain current location information. The input is the signal from the GPS module and the output is location data. The device uses the GPS module to obtain real-time location information (latitude and longitude) and saves the data in the application.

[1107] Step 4:

[1108] The device collects video data, audio data, and location information and combines them into packets. The input is each file and location data, and the output is a data packet. The device combines the video file, audio file, and location information into a single packet.

[1109] Step 5:

[1110] The terminal sends a data packet to the server. The input is the data packet, and the output is a notification of successful data transmission to the server. The terminal sends the data packet to the server via the Internet and confirms its success.

[1111] Step 6:

[1112] The server analyzes the received data. The input is a data packet, and the output is the emergency identification result. The server separates video, audio, and location information from the received data packet, analyzes this data using a generative AI model, and identifies the type of emergency.

[1113] Step 7:

[1114] The server determines the most appropriate emergency contact based on the analysis results. The input is the analysis results, and the output is emergency contact information. Based on the analysis results, the server selects the most appropriate emergency service (ambulance, police, fire department, etc.).

[1115] Step 8:

[1116] The device receives instructions from the server and automatically calls the appropriate emergency service. The input is emergency contact information, and the output is a successful call. The device automatically calls the emergency contact provided by the server and starts the call.

[1117] Step 9:

[1118] The device automatically explains the situation using a voice message generated by the server. The input is the voice message from the server, and the output is the completion of sending the voice message. The device plays the voice message generated by the server and explains the situation to the emergency services.

[1119] Step 10:

[1120] The terminal provides the user with a first aid guide prepared in advance on the server side. The input is first aid guide information, and the output is completion of guide provision. The terminal provides the first aid guide provided by the server to the user in voice or text format.

[1121] Step 11:

[1122] The server tracks the user's real-time location and calculates the optimal route to the nearest hospital or fire station. The input is the user's current location information and the output is the optimal route information. The server uses GPS data to identify the user's current location and calculates the optimal route to the nearest emergency facility.

[1123] Step 12:

[1124] The terminal provides optimal route navigation to the user. The input is optimal route information, and the output is navigation completion. The terminal provides turn-by-turn navigation to the user based on the optimal route information provided by the server.

[1125] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1126] The present invention is a system that takes into consideration the user's emotions in an emergency and can respond quickly and appropriately. This system provides a means for responding to an emergency through a smartphone app, and provides appropriate guidance according to the user's stress and impatience.

[1127] System configuration

[1128] This system consists of the user's smartphone device (hereafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. The server runs a multimodal generation AI and an emotion recognition engine, which analyzes the transmitted data.

[1129] System Operation

[1130] 1. Pressing the emergency button and collecting data

[1131] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously enables the microphone to record audio, acquires location information using the GPS module, and packetizes the audio data for transmission to the server.

[1132] Example: When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and obtaining location information.

[1133] 2. Data submission and analysis

[1134] The device then combines the captured video and audio files, along with location information, into a single data packet and sends it to the server, which then queues the data for analysis and activates the multimodal generation AI and emotion recognition engine.

[1135] Example: The server analyzes the data it receives, identifies traffic accidents and emergencies from voices, and recognizes emotions from the user's tone of voice.

[1136] 3. Emotion Recognition and Emergency Service Decisions

[1137] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as anxiety, fear, panic, etc. The server then uses the analysis results to determine the type and severity of the emergency and selects the appropriate emergency services (e.g., police, ambulance, fire department, etc.).

[1138] Example: If the server detects that the user is panicking, it will determine the level of urgency as high and send instructions to the device to immediately contact an ambulance and the police.

[1139] 4. Automatic contact and first aid provision

[1140] The device receives instructions from the server and automatically calls the selected emergency service contact. The device plays a voice message generated by the server, explaining the situation. Furthermore, the device adjusts the first aid guidance content according to the user's emotional state and provides it to the user.

[1141] Example: For a user in a panicked state, a calming voice will provide instructions on CPR and how to treat injuries.

[1142] 5. Providing location information and navigation

[1143] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information, and the device displays this information to the user and provides turn-by-turn navigation.

[1144] Examples include providing detailed directions to the nearest hospital to help users avoid getting lost, and continuing to provide first aid instructions even when communication is unstable, using pre-installed language models.

[1145] 6. Providing Additional Information

[1146] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions, so that emergency services can make appropriate preparations before arriving on scene.

[1147] Example: If the user is very disoriented, this information can also be communicated to emergency responders so that they can provide psychological support when they arrive on scene.

[1148] In this way, the system, including the emotion recognition engine, enables a fast and accurate response in emergency situations, which is expected to increase the user's sense of security and improve the overall survival rate.

[1149] The processing flow will be explained below.

[1150] Step 1:

[1151] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously activates the microphone to capture surrounding audio, and uses the GPS module to obtain the device's current location.

[1152] Step 2:

[1153] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[1154] Step 3:

[1155] The server adds the received data packets to an analysis queue and activates a multimodal generation AI and emotion recognition engine, which analyzes video data to identify emergency situations and recognizes user emotions from audio data.

[1156] Step 4:

[1157] The server's emotion recognition engine analyzes the user's tone of voice, speed, and word choice from the audio data to determine whether the user is anxious, calm, or panicked.

[1158] Step 5:

[1159] Based on the analysis results, the server determines the type and severity of the emergency, selects the most appropriate emergency service contact, and sends this information back to the device.

[1160] Step 6:

[1161] Based on the instructions received, the device automatically calls the selected emergency service contact and plays a server-generated voice message explaining the situation.

[1162] Step 7:

[1163] The device provides first aid instructions to the user. If the server's emotion recognition engine determines that the user's emotional state is impatient or panicked, the device adjusts the content and speed of the instructions to help the user calm down.

[1164] Step 8:

[1165] The server calculates the optimal route to the nearest hospital or fire station based on real-time location information and sends that information to the device, which then displays the calculated route information to the user and begins turn-by-turn navigation.

[1166] Step 9:

[1167] The emotion recognition engine takes into account the user's emotions and provides additional information to emergency contacts, allowing them to make appropriate preparations before emergency services arrive on scene.

[1168] Step 10:

[1169] By following the instructions on the device and receiving first aid and navigation, users can take appropriate action until emergency services arrive, which will enable quick and accurate response and is expected to improve the survival rate.

[1170] Example 2

[1171] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1172] Conventional emergency response systems are required to respond not only based on user reports but also by taking into account the user's emotional state. However, current systems have difficulty accurately recognizing the user's emotions and providing appropriate guidance and emergency measures. Another challenge is providing accurate responses in environments with unstable communication conditions.

[1173] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1174] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for adjusting first aid guidance according to the user's emotional state, and a means for calculating an optimal route based on the user's real-time location information and providing navigation, thereby enabling a quick and accurate emergency response that also takes the user's emotional state into consideration.

[1175] 1. "Image capture means" refers to a device that activates a camera and captures video or still images when a user presses a trigger.

[1176] 2. "Location information acquisition means" refers to a device that acquires location information of the user's current location using a GPS module or the like.

[1177] 3. "Transmitting means" refers to a device that packetizes the acquired image data, audio data, and location information and transmits them to the server.

[1178] 4. "Analysis means" is a device that analyzes the data transmitted by the transmission means and identifies an emergency situation or the user's emotions.

[1179] 5. "Contact Means" means a device that determines the appropriate emergency contact based on the emergency situation identified by the Analysis Means and automatically initiates a call.

[1180] 6. "Guidance means" is a device that provides the user with first aid guidance and appropriate actions according to their emotional state.

[1181] 7. "Means for providing navigation" means a device that calculates the optimal route based on the user's real-time location information and provides turn-by-turn navigation.

[1182] 8. "Acoustic analysis means" refers to a device that analyzes the surrounding acoustic environment and the user's emotions based on audio data.

[1183] 9. "Language model" refers to a model that is installed to function even in environments with unstable communication and is used to generate emergency and guidance messages.

[1184] MODE FOR CARRYING OUT THE INVENTION

[1185] This invention is an emergency response system that takes into account the user's emotions and can respond quickly and appropriately in an emergency. This system operates via a smartphone app and handles emergencies using an emotion recognition engine and multimodal generation AI.

[1186] Hardware Configuration

[1187] The hardware configuration of this system is as follows:

[1188] 1. Device: A smartphone carried by the user, equipped with the following functions:

[1189] Camera features

[1190] Microphone function

[1191] GPS Modules

[1192] Internet communication function

[1193] 2. Server: A back-end system connected to terminals via the Internet, and includes the following components:

[1194] Multimodal generation AI (audio data analysis, image data analysis)

[1195] Emotion Recognition Engine

[1196] Data Analysis Engine

[1197] Software Configuration

[1198] 1. Terminal application: A smartphone application used by the user that performs the following actions in the event of an emergency.

[1199] Emergency button activates camera and microphone

[1200] Obtaining location information

[1201] Packetizing acquired data (video, audio, location information) and sending it to the server

[1202] 2. Server application: Software that runs on the server side and provides the following functions:

[1203] Analyzing received data

[1204] Recognizing the user's emotional state

[1205] Determining the type and severity of an emergency

[1206] Selection of emergency contacts and automatic call instructions

[1207] Generate and submit a first aid guide

[1208] Real-time location-based navigation

[1209] Example of operation

[1210] To give a concrete example of how this works, consider the following scenario:

[1211] 1. If the emergency button is pressed:

[1212] A user encounters a traffic accident and presses the emergency button on their smartphone.

[1213] The device immediately activates the camera and records video of the accident scene.

[1214] At the same time, the microphone is enabled to record audio and location information is obtained using GPS.

[1215] The terminal transmits this data to the server.

[1216] 2. Data transmission and analysis:

[1217] The server adds the received video, audio, and location information to an analysis queue and activates the multimodal generation AI and emotion recognition engine.

[1218] The server analyzes the user's emotional state (e.g., panic) from the tone of the voice and determines the appropriate emergency response.

[1219] 3. Contacting the appropriate emergency services:

[1220] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and issues an automatic call instruction.

[1221] The device automatically contacts emergency services and plays a server-generated voice message explaining the situation.

[1222] Prompt Sentence Examples

[1223] Examples of prompts include:

[1224] "A user encounters a traffic accident and presses the emergency button. The mobile device collects video, audio, and GPS data and sends it to a server. The server analyzes the urgency of the accident and coordinates the appropriate emergency services. Please explain how the system works."

[1225] This system enables a fast and effective response in emergency situations while taking into consideration the user's feelings, which is expected to improve the user's sense of security and increase the overall survival rate.

[1226] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1227] Step 1:

[1228] The user presses the emergency button.

[1229] Input: The user presses the emergency button.

[1230] Output: The emergency button press signal is transmitted to the terminal.

[1231] The device immediately activates the camera and starts recording video, while simultaneously activating the microphone for audio recording and the GPS module for location information.

[1232] Specific operation: Video recording, audio recording, and location information acquisition of the accident scene are carried out simultaneously.

[1233] Step 2:

[1234] The terminal combines the acquired video file, audio file, and location information into a single data packet and sends it to the server.

[1235] Input: Video data acquired by the camera, audio data recorded by the microphone, and location information acquired by GPS.

[1236] Output: A data packet is generated and sent to the server.

[1237] Specific operation: Various data is packetized and sent to a server via the Internet.

[1238] Step 3:

[1239] The server adds the received data to a queue for analysis and processing, and activates the multimodal generation AI and emotion recognition engine.

[1240] Input: Data packets sent from the terminal.

[1241] Output: The analysis results include the type of emergency and the user's emotional state.

[1242] Specific operation: The received data is added to the analysis queue, and the AI ​​and emotion recognition engine are activated to begin data analysis.

[1243] Step 4:

[1244] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as impatience, fear, or panic.

[1245] Input: Audio data.

[1246] Output: Analysis results showing the user's emotional state.

[1247] Specific operation: Emotion analysis is performed using voice data to determine the user's emotional state.

[1248] Step 5:

[1249] Based on the analysis results, the server determines the type and urgency of the emergency and selects the appropriate emergency service (e.g., police, ambulance, fire department, etc.).

[1250] Input: The type of emergency situation as the analysis result and the user's emotional state.

[1251] Output: Instructions for optimal emergency service selection.

[1252] Specific operation: Based on the analysis results, the level of urgency is determined to be high, and appropriate emergency services are immediately selected.

[1253] Step 6:

[1254] The device receives instructions from the server and automatically calls the selected emergency service contact.

[1255] Input: Instructions from the server.

[1256] Output: Initiation of automatic contact with emergency services.

[1257] Specific operation: The device makes an automatic call and plays a voice message generated by the server to explain the situation.

[1258] Step 7:

[1259] The device provides first aid guidance according to the user's emotional state.

[1260] Input: Analysis results indicating the user's emotional state.

[1261] Output: First aid guide depending on emotional state.

[1262] Specific actions: Provide breathing techniques and first aid instructions in a calm voice.

[1263] Step 8:

[1264] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information.

[1265] Input: The user's real-time location.

[1266] Output: Optimal route information.

[1267] Specific operation: Analyze real-time location information and calculate the optimal route.

[1268] Step 9:

[1269] The device displays optimal route information to the user and provides turn-by-turn navigation.

[1270] Input: Optimal route information.

[1271] Output: Navigation guide.

[1272] Specific behavior: Support the user by providing detailed directions to the nearest hospital.

[1273] Step 10:

[1274] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions.

[1275] Input: Data indicating the user's emotional state.

[1276] Output: Additional information for emergency contacts.

[1277] Specific behavior: If the user is confused, this information will also be conveyed to emergency responders so that psychological support can be prepared when they arrive at the scene.

[1278] (Application example 2)

[1279] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1280] In the event of an emergency, if a user feels anxious or scared, it can be difficult to respond appropriately. Furthermore, conventional emergency response systems do not take the user's emotional state into account, which can lead to inadequate responses. Furthermore, the system must function even in environments with unstable communications. The present invention aims to solve these problems and provide a system that allows users to respond to emergencies with peace of mind.

[1281] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1282] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, an emotion recognition means for analyzing the user's emotional state identified by the analysis means, a contact means for determining an appropriate emergency contact based on the analysis result of the emotion recognition means and automatically initiating a call, and a guidance means for providing guidance on appropriate first aid according to the user's emotional state. This enables a quick and appropriate emergency response while taking the user's emotional state into consideration.

[1283] "Image capture means" refers to a technical means for capturing an image or video using a device such as a camera when a user activates a trigger.

[1284] "Location information acquisition means" refers to a technical means for determining the current location of a terminal using GPS or other positioning systems.

[1285] The "transmission means" is a means having a function of transmitting the acquired image data, audio data, and location information to a server or other external system.

[1286] "Analysis Means" means the technical means used to identify an emergency situation and determine appropriate emergency contacts based on the Transmitted Data.

[1287] "Emotion recognition means" is a technical means for analyzing the user's emotional state from voice data, etc., and identifying emotions such as impatience or fear.

[1288] A "contact method" is a method that has the ability to automatically initiate a call to an emergency contact based on an identified emergency condition and emotional state.

[1289] The "guidance means" is a technical means for providing a first aid guide in response to an emergency situation faced by a user, taking into consideration the user's emotional state.

[1290] MODE FOR CARRYING OUT THE INVENTION

[1291] To implement this invention, the user's smartphone terminal, the server, and the emergency service work together. The hardware and software used in each step and the processing content thereof will be specifically described below.

[1292] System configuration

[1293] 1. Hardware

[1294] Device (smartphone): Smartphones are equipped with a camera, microphone, and GPS module. These functions are used to acquire images, audio, and location information.

[1295] Server: The server runs a multimodal generation AI and emotion recognition engine, which analyzes the transmitted data.

[1296] 2. Software

[1297] Image capture tool: Software that activates the camera and captures images and videos.

[1298] Audio recording software: Records audio using a microphone and generates audio data.

[1299] Location information acquisition software: Uses the GPS module to determine the device's current location.

[1300] Data transmission software: Collects the acquired data into a single data packet and sends it to a server over the Internet.

[1301] Emotion recognition engine: Analyzes the user's emotional state from voice data and identifies emotions such as impatience and fear. Specifically, it uses an emotion recognition model from the Transformers library.

[1302] Multimodal generative AI: Identifies emergencies based on submitted data and generates appropriate responses. Uses the nlptown / bert-base-multilingual-uncased-sentiment model.

[1303] Contact Method: Automatically call emergency contacts based on the analyzed results.

[1304] Guidance: Provides first aid guidance according to the user's emotional state. Even if communication is unstable, the guidance continues using a pre-installed language model.

[1305] System Operation

[1306] 1. Pressing the emergency button and collecting data

[1307] When a user presses the emergency button, the device's camera automatically activates and begins recording images and videos. At the same time, the microphone activates and records audio. The GPS module also activates and collects location information. This data is packetized within the device and sent to the server.

[1308] 2. Data Analysis and Emotion Recognition

[1309] The server analyzes the received data and identifies the user's emotional state from the user's voice tone and environmental sounds. An emotion recognition engine analyzes the voice data and identifies whether the user is expressing emotions such as anxiety or fear, and determines the appropriate emergency measures.

[1310] Examples:

[1311] When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and acquiring location information. The server analyzes this data to understand the circumstances of the accident and the user's emotional state.

[1312] 3. Selecting and contacting emergency services

[1313] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and automatically calls the contacts. The server generates a voice message explaining the situation.

[1314] 4. First aid information

[1315] The system takes into account the user's emotional state and provides the most appropriate first aid guidance. For example, if a user is in a panic, a slow voice will guide them through CPR and how to treat injuries.

[1316] Examples:

[1317] By inputting the following prompts into the generative AI model, specific emergency response instructions can be generated.

[1318] Example prompt:

[1319] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[1320] As a result, it is possible to take prompt and appropriate emergency measures while taking into consideration the emotional state of the user.

[1321] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1322] Step 1:

[1323] When a user presses the emergency button, the device immediately activates the camera and starts recording images and videos. At the same time, it activates the microphone to record audio and uses the GPS module to obtain location information. These data are packetized and ready to be transmitted. The input is the user pressing the button, and the output is image data, audio data, and location information.

[1324] Step 2:

[1325] The device collects the acquired image data, audio data, and location information into a single data packet and sends it to a server via the Internet. The input is the image data, audio data, and location information, and the output is the data packet sent to the server.

[1326] Step 3:

[1327] The server analyzes the received data packets using a multimodal generation AI and emotion recognition engine. First, it identifies the emergency situation from image data and location information, and then analyzes the audio data to recognize the user's emotional state. The input is the data packets, and the output is the emergency situation identification result and the user's emotional state recognition result.

[1328] Step 4:

[1329] The server analyzes the voice data using an emotion recognition engine to determine whether the user is expressing emotions such as anxiety or fear. For example, it uses an emotion recognition model from the transformers library. The input is the voice data, and the output is the identification of the emotional state.

[1330] Step 5:

[1331] The server determines the appropriate emergency contact based on the analysis results and instructs the device to automatically initiate a call. The server generates the necessary call content using a generative AI model and sends it to the device. The input is the emergency state and emotional state identification results, and the output is the emergency contact and call content.

[1332] Step 6:

[1333] The device receives instructions from the server and automatically calls the appropriate emergency contact. It responds by playing a voice message generated by the server and explaining the situation. The input is the call instruction from the server, and the output is connecting the phone and explaining the situation to the other party.

[1334] Step 7:

[1335] The server uses a generative AI model to create a first aid guide based on the user's emotional state and sends it to the device. The device then provides this guide to the user. Even if communication is unstable, the device continues providing guidance using the installed language model. The input is the user's emotional state and situation data, and the output is the first aid guide content.

[1336] Specific examples

[1337] The following prompts can be fed into a generative AI model to generate specific emergency response instructions:

[1338] Example prompt:

[1339] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[1340] The above series of steps allows for a quick and appropriate emergency response while taking into account the user's emotional state.

[1341] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1342] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1343] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1344] [Fourth embodiment]

[1345] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1346] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1347] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1348] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1349] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1350] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1351] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1352] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1353] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1354] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1355] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1356] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1357] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1358] The present invention is a system for responding quickly and appropriately to emergencies. The system is built around a smartphone application that allows users who encounter an emergency to automatically contact the appropriate emergency services with the push of a button.

[1359] System configuration

[1360] This system consists of a user's smartphone device (hereinafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. A multimodal generation AI runs on the server side and analyzes the transmitted data.

[1361] System Operation

[1362] 1. Pressing the emergency button and collecting data

[1363] When a user presses the emergency button, the device activates the camera and starts recording video from that point onward, simultaneously recording audio data using the microphone, and obtaining the device's current location using the GPS module.

[1364] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[1365] 2. Data submission and analysis

[1366] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[1367] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[1368] 3. Determine emergency services and initiate a call

[1369] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[1370] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[1371] 4. First Aid and Navigation

[1372] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station, allowing the device to provide turn-by-turn navigation to the user.

[1373] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[1374] This system will enable users to respond quickly and accurately in emergency situations, which is expected to improve the chances of survival. In addition, by utilizing speech analysis methods and language models that work even in unstable communication environments, appropriate support can be provided in any situation.

[1375] The processing flow will be explained below.

[1376] Step 1:

[1377] When a user presses the emergency button, the device immediately activates the camera and starts recording video. At the same time, the microphone is activated to record surrounding audio. The device also uses the GPS module to obtain its current location.

[1378] Step 2:

[1379] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[1380] Step 3:

[1381] The server adds the received data packets to the analysis queue and activates the multimodal generation AI. The server analyzes the video data to identify emergency situations (e.g., traffic accidents, fires, collapsed people, etc.) and analyzes the surrounding sound environment from the audio data.

[1382] Step 4:

[1383] Based on the analysis results, the server determines the type and severity of the emergency, identifies the current geographic location from the location information, and lists the nearest emergency services (e.g., police station, fire station, hospital, etc.).

[1384] Step 5:

[1385] The server selects the most appropriate emergency service contact based on the determined emergency situation and returns that information to the device.

[1386] Step 6:

[1387] The device responds to the received instructions and automatically calls the selected emergency service contact, and plays a voice message generated by the server explaining the situation.

[1388] Step 7:

[1389] After the call is made, the device will begin providing first aid instructions to the user. The server will calculate the optimal route to the nearest hospital or fire station based on the user's location and send that information to the device.

[1390] Step 8:

[1391] The device will display the calculated optimal route information to the user, initiate turn-by-turn navigation, and even if the connection is unstable, it can continue to provide first aid instructions using pre-installed language models.

[1392] Step 9:

[1393] By following the first aid and navigation instructions, users can facilitate appropriate responses at the scene until emergency services arrive, which can lead to faster and more accurate responses and potentially improved survival rates.

[1394] Example 1

[1395] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1396] When encountering an emergency, it can be difficult for individuals to contact emergency services quickly and appropriately. This is especially true when users are in a state of panic or have difficulty explaining the situation, which can delay appropriate responses. Furthermore, there is a risk that damage may escalate due to users not knowing how to respond appropriately. There is a need for a system that can resolve these issues and provide a fast and appropriate emergency response.

[1397] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1398] In this invention, the server includes means for activating an image acquisition means and a voice acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, voice data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for calculating an optimal route based on the user's location information and providing navigation, thereby enabling the user to respond to an emergency quickly and accurately.

[1399] "Image acquisition means" refers to a device or function that uses a camera or other image sensor to acquire video or image information about the surroundings.

[1400] An "audio capture means" is a device or function that uses a microphone or other audio sensor to record or capture ambient sounds.

[1401] "Location Information Acquisition Means" means a device or function that acquires current geographic location information using GPS or other location determination technology.

[1402] The "transmission means" is a function that assembles the acquired data into a single packet and transmits it via the Internet or other communication means.

[1403] "Analysis means" refers to software or hardware that analyzes received data and identifies specific conditions or information.

[1404] "Multimodal generation AI" is an artificial intelligence technology that integrates and analyzes multiple modals (input formats) such as images, audio, and text.

[1405] "Emergency Contacts" are contact details for agencies and services (e.g., ambulance, police, fire department, etc.) that will best respond to a particular emergency.

[1406] "Contact methods" are devices or functions that automatically initiate a call to an emergency contact and convey the necessary information.

[1407] The "guidance means" is a function for providing first aid and navigation information to the user.

[1408] "Navigation means" is a function that calculates the optimal route based on the user's location information and provides real-time directions.

[1409] MODE FOR CARRYING OUT THE INVENTION

[1410] The present invention is a system for responding promptly and appropriately to an emergency situation, and functions in cooperation with a user, a terminal, and a server. A specific embodiment of this system will be described below.

[1411] 1. Components and Hardware Configuration

[1412] Terminal

[1413] The device is a smartphone owned by the user and is equipped with the following hardware:

[1414] Camera: Captures video and images of the surrounding area.

[1415] Microphone: Records surrounding sounds.

[1416] GPS module: Obtains current location information.

[1417] These hardware components function as an "image acquisition means," "audio acquisition means," and "location information acquisition means." In addition, a dedicated application is installed, and these functions are automatically activated when the emergency button is pressed.

[1418] server

[1419] The server will be deployed in a cloud environment and will utilize the following software and technologies:

[1420] Multimodal generative AI: Integrated analysis of video, audio, and location data to identify emergencies.

[1421] Communication means: Receives data sent from the terminal and sends the analysis results to the terminal.

[1422] 2. Usage and Data Processing

[1423] Pressing the emergency button

[1424] When a user presses the emergency button within a smartphone application, the device automatically activates the camera, microphone, and GPS to collect data.

[1425] Sending data

[1426] The terminal combines the collected video data, audio data, and location information into a single data packet, which is then transmitted to a server via the Internet by a transmitting means.

[1427] Data analysis

[1428] The server analyzes the received data packets using an analytical method (multimodal generative AI), which identifies the type of emergency and selects the appropriate emergency services (ambulance, police, fire department, etc.).

[1429] Example: If the server analyzes the data it receives and determines that it is a "traffic accident" based on video and audio, the analysis results will also include the extent of injuries and the danger level of the scene.

[1430] Starting a call

[1431] The device receives instructions from the server and automatically calls the designated emergency service, explaining the situation using a voice message generated by the server.

[1432] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, providing details of the emergency situation via voice.

[1433] First Aid Guide and Navigation

[1434] While waiting for the ambulance to arrive, the device provides a pre-prepared first aid guide from the server. The server also tracks the user's location in real time and calculates the optimal route to the nearest hospital or fire station. As a navigation tool, the device provides turn-by-turn navigation information to the user.

[1435] Example: The device provides voice and text instructions for CPR and displays directions to the nearest hospital.

[1436] Prompt Sentence Examples

[1437] An example of a prompt to be input to the generative AI model when a user encounters an emergency:

[1438] "I was involved in a traffic accident and someone was injured. My current location is latitude xx.xxxx, longitude yy.yyyy."

[1439] "A fire has broken out and there are many injured people nearby. Your current location is latitude xx.xxxx, longitude yy.yyyy."

[1440] In this way, the entire system works together to provide comprehensive support in emergencies, enabling users to quickly take the most appropriate action.

[1441] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1442] Step 1:

[1443] User operations

[1444] The user presses the emergency button in the application.

[1445] Input: User touch input.

[1446] Output: Emergency button pressed event.

[1447] Specific action: The user taps the emergency button displayed on the smartphone screen.

[1448] Step 2:

[1449] Device behavior: Data collection

[1450] The device detects when the emergency button is pressed and activates the camera, microphone, and GPS module to collect data.

[1451] Input: Emergency button press event.

[1452] Output: Video data, audio data, location information.

[1453] Specific behavior:

[1454] The camera will activate and record video of the emergency.

[1455] The microphone records the audio.

[1456] The GPS module obtains the current location information.

[1457] Step 3:

[1458] Device operation: Data transmission

[1459] The device bundles the collected data (video data, audio data, location information) into a single data packet and sends it to the server.

[1460] Input: Video data, audio data, location information.

[1461] Output: Data packets sent to the server.

[1462] Specific behavior:

[1463] Video, audio, and location information are packaged into packets.

[1464] The packet is sent over the Internet to a server.

[1465] Step 4:

[1466] Server Operation: Data Analysis

[1467] The server analyzes the received data packets and runs a multimodal generative AI to identify emergencies.

[1468] Input: Data packet.

[1469] Output: Emergency type, detailed analysis results.

[1470] Specific behavior:

[1471] The received video data, audio data, and location information are expanded.

[1472] Multimodal generative AI analyzes the data, identifies the type of emergency (e.g., traffic accident, fire, sudden illness, etc.), and determines the detailed situation.

[1473] Step 5:

[1474] Server Behavior: Emergency Services Decision

[1475] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, or fire department).

[1476] Input: Emergency type, detailed analysis results.

[1477] Output: Emergency services contact information.

[1478] Specific behavior:

[1479] Retrieve the best contact information from a database for the type of emergency.

[1480] Send contact information for selected emergency services to the device.

[1481] Step 6:

[1482] Device Action: Initiating a call

[1483] The device receives instructions from the server and automatically calls emergency services. During the call, it plays a voice message generated by the server explaining the situation.

[1484] Input: Emergency services contact information, server-generated voice message.

[1485] Output: Calls to emergency services, playback of voice messages.

[1486] Specific behavior:

[1487] Your device will automatically dial the emergency services number.

[1488] Once the call is connected, a server-generated voice message is played, detailing the emergency.

[1489] Step 7:

[1490] Device Operation: First Aid and Navigation

[1491] While waiting for the ambulance to arrive, the device will provide the user with first aid guidance and navigate the optimal route to the nearest hospital or fire station.

[1492] Input: First aid guide information provided by the server, real-time location information.

[1493] Output: First aid guide, navigation information.

[1494] Specific behavior:

[1495] The device provides the user with pre-programmed first aid instructions (audio, text, video).

[1496] The server calculates the optimal route based on the user's location information and sends navigation information to the terminal.

[1497] The device displays real-time turn-by-turn navigation to the user.

[1498] (Application example 1)

[1499] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1500] Conventional emergency response systems have the drawback of making it difficult for users to respond quickly and appropriately when they encounter an emergency. In particular, they require complex procedures, such as initiating a call to an emergency contact and providing first aid instructions, which often prevents users from responding calmly. Furthermore, they lack the functionality to provide an optimal route based on the user's current location, making it difficult to evacuate to a safe location efficiently. Therefore, there is a need for a system that supports users' quick and accurate actions in emergency response.

[1501] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1502] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for providing the user with first aid guidance, and a means for the guidance means to calculate an optimal travel route based on the user's current location and provide navigation, thereby enabling the user to respond quickly and accurately to an emergency situation and efficiently evacuate to a safe place.

[1503] - A "trigger" is an input device such as a button or switch that a user operates to initiate a specific action.

[1504] "Image acquisition means" refers to a device that collects video data using a camera or the like.

[1505] "Location information acquisition means" is a device that identifies the current location using a GPS module or the like.

[1506] "Audio data" refers to a recording of sound captured using a microphone or the like.

[1507] "Transmission means" means a device or process that transmits collected data to a server via the Internet or other communication means.

[1508] "Analysis means" refers to algorithms or software that process received data and identify emergency conditions.

[1509] A "call" is a means of making a voice communication, which may be made over a telephone line or the Internet.

[1510] "Contact Assistance" means any device or software that automatically places a call to the appropriate emergency service.

[1511] A "guiding means" is a device or software that provides useful information to a user.

[1512] A "First Aid Guide" is information that provides instructions on how to provide basic medical assistance in an emergency.

[1513] "Navigation" is a route guidance function for guiding a user to a specific destination.

[1514] "Current location of user" is real-time location information of the user identified by the location information acquisition means.

[1515] This invention is a system for responding quickly and accurately to emergencies, and is built around a user's smartphone terminal. The configuration and operation of the system are described in detail below.

[1516] 1. System Configuration

[1517] This system consists of a user's smartphone terminal (hereinafter referred to as the terminal), a server connected via the Internet, and an emergency service. The terminal is equipped with a camera, microphone, GPS, and the application of this invention. A multimodal generation AI runs on the server side.

[1518] 2. System Operation

[1519] Pressing the emergency button and collecting data

[1520] When the user presses the emergency button, the application installed on the device activates the camera and starts recording video from that point. At the same time, it also records audio data using the microphone. It also obtains the current location information using the GPS module.

[1521] Example: If a user is at the scene of a traffic accident, pressing the emergency button will cause the device to take a video of the scene, record the sound of the accident, and obtain location information.

[1522] Data transmission and analysis

[1523] The device then collects video, audio, and location data and sends it to a server via the internet. The server then instantly analyzes the data and activates a multimodal generative AI to identify the type of emergency.

[1524] Example: The server analyzes the received data and determines that it is a "traffic accident" based on video and audio from the accident scene, and the analysis results also include the extent of injuries and the degree of danger to the surrounding area.

[1525] Determining emergency services and initiating a call

[1526] Based on the analysis results, the server determines the most appropriate emergency service (e.g., ambulance, police, fire department, etc.). The device receives this instruction and automatically calls the appropriate emergency service on behalf of the user. The device also automatically explains the situation using a voice message generated by the server.

[1527] Example: In the case of a traffic accident, the device will automatically call an ambulance and the police, and provide a voice message with the location of the accident and the status of any injuries.

[1528] First Aid and Navigation

[1529] While waiting for the ambulance to arrive, the device will provide the user with a pre-prepared first aid guide, and the server will track the user's location in real time and calculate the optimal route to the nearest hospital or fire station based on the user's current location, allowing the device to provide turn-by-turn navigation to the user.

[1530] Example: The device provides voice and text instructions to the user on how to perform CPR and also provides directions to the nearest hospital.

[1531] Prompt Sentence Examples

[1532] This is an emergency. A user has suddenly lost consciousness at home. Location, video, and audio data are attached below. Please take appropriate action.

[1533] Location: Tokyo, Japan

[1534] Video data: emergency_video.avi

[1535] Audio data: emergency_audio.wav

[1536] Hardware and software used

[1537] Device: Smartphone (camera, microphone, GPS)

[1538] Server: Internet Server

[1539] Software: Multimodal generative AI, Python, OpenCV, PyAudio, Geopy, Requests library

[1540] This allows users to respond quickly and accurately to emergencies and efficiently evacuate to a safe location. In addition, the server tracks users' locations in real time and provides first aid guidance and optimal route navigation, which is expected to improve survival rates.

[1541] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1542] Step 1:

[1543] The device detects the pressing of the emergency button. The input is the user pressing the button, and the output is the activation of the camera, microphone, and GPS. When the user performs a trigger operation, the device detects this event and activates the camera, microphone, and GPS sensors.

[1544] Step 2:

[1545] The device records video with the camera and audio with the microphone. The input is data information from the camera and microphone, and the output is a video file and an audio file. The device starts the camera, records video for a certain period of time (e.g., 10 seconds), and simultaneously records audio data using the microphone. These data are then saved to a file.

[1546] Step 3:

[1547] The device uses GPS to obtain current location information. The input is the signal from the GPS module and the output is location data. The device uses the GPS module to obtain real-time location information (latitude and longitude) and saves the data in the application.

[1548] Step 4:

[1549] The device collects video data, audio data, and location information and combines them into packets. The input is each file and location data, and the output is a data packet. The device combines the video file, audio file, and location information into a single packet.

[1550] Step 5:

[1551] The terminal sends a data packet to the server. The input is the data packet, and the output is a notification of successful data transmission to the server. The terminal sends the data packet to the server via the Internet and confirms its success.

[1552] Step 6:

[1553] The server analyzes the received data. The input is a data packet, and the output is the emergency identification result. The server separates video, audio, and location information from the received data packet, analyzes this data using a generative AI model, and identifies the type of emergency.

[1554] Step 7:

[1555] The server determines the most appropriate emergency contact based on the analysis results. The input is the analysis results, and the output is emergency contact information. Based on the analysis results, the server selects the most appropriate emergency service (ambulance, police, fire department, etc.).

[1556] Step 8:

[1557] The device receives instructions from the server and automatically calls the appropriate emergency service. The input is emergency contact information, and the output is a successful call. The device automatically calls the emergency contact provided by the server and starts the call.

[1558] Step 9:

[1559] The device automatically explains the situation using a voice message generated by the server. The input is the voice message from the server, and the output is the completion of sending the voice message. The device plays the voice message generated by the server and explains the situation to the emergency services.

[1560] Step 10:

[1561] The terminal provides the user with a first aid guide prepared in advance on the server side. The input is first aid guide information, and the output is completion of guide provision. The terminal provides the first aid guide provided by the server to the user in voice or text format.

[1562] Step 11:

[1563] The server tracks the user's real-time location and calculates the optimal route to the nearest hospital or fire station. The input is the user's current location information and the output is the optimal route information. The server uses GPS data to identify the user's current location and calculates the optimal route to the nearest emergency facility.

[1564] Step 12:

[1565] The terminal provides optimal route navigation to the user. The input is optimal route information, and the output is navigation completion. The terminal provides turn-by-turn navigation to the user based on the optimal route information provided by the server.

[1566] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1567] The present invention is a system that takes into consideration the user's emotions in an emergency and can respond quickly and appropriately. This system provides a means for responding to an emergency through a smartphone app, and provides appropriate guidance according to the user's stress and impatience.

[1568] System configuration

[1569] This system consists of the user's smartphone device (hereafter referred to as the device), a server connected via the Internet, and an emergency service. The device has a camera, microphone, GPS, and an application installed. The server runs a multimodal generation AI and an emotion recognition engine, which analyzes the transmitted data.

[1570] System Operation

[1571] 1. Pressing the emergency button and collecting data

[1572] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously enables the microphone to record audio, acquires location information using the GPS module, and packetizes the audio data for transmission to the server.

[1573] Example: When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and obtaining location information.

[1574] 2. Data submission and analysis

[1575] The device then combines the captured video and audio files, along with location information, into a single data packet and sends it to the server, which then queues the data for analysis and activates the multimodal generation AI and emotion recognition engine.

[1576] Example: The server analyzes the data it receives, identifies traffic accidents and emergencies from voices, and recognizes emotions from the user's tone of voice.

[1577] 3. Emotion Recognition and Emergency Service Decisions

[1578] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as anxiety, fear, panic, etc. The server then uses the analysis results to determine the type and severity of the emergency and selects the appropriate emergency services (e.g., police, ambulance, fire department, etc.).

[1579] Example: If the server detects that the user is panicking, it will determine the level of urgency as high and send instructions to the device to immediately contact an ambulance and the police.

[1580] 4. Automatic contact and first aid provision

[1581] The device receives instructions from the server and automatically calls the selected emergency service contact. The device plays a voice message generated by the server, explaining the situation. Furthermore, the device adjusts the first aid guidance content according to the user's emotional state and provides it to the user.

[1582] Example: For a user in a panicked state, a calming voice will provide instructions on CPR and how to treat injuries.

[1583] 5. Providing location information and navigation

[1584] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information, and the device displays this information to the user and provides turn-by-turn navigation.

[1585] Examples include providing detailed directions to the nearest hospital to help users avoid getting lost, and continuing to provide first aid instructions even when communication is unstable, using pre-installed language models.

[1586] 6. Providing Additional Information

[1587] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions, so that emergency services can make appropriate preparations before arriving on scene.

[1588] Example: If the user is very disoriented, this information can also be communicated to emergency responders so that they can provide psychological support when they arrive on scene.

[1589] In this way, the system, including the emotion recognition engine, enables a fast and accurate response in emergency situations, which is expected to increase the user's sense of security and improve the overall survival rate.

[1590] The processing flow will be explained below.

[1591] Step 1:

[1592] When a user presses the emergency button, the device immediately activates the camera and starts recording video, simultaneously activates the microphone to capture surrounding audio, and uses the GPS module to obtain the device's current location.

[1593] Step 2:

[1594] The device compiles the acquired video files, audio files, and location information to generate a data packet, which is then sent to a server via the Internet.

[1595] Step 3:

[1596] The server adds the received data packets to an analysis queue and activates a multimodal generation AI and emotion recognition engine, which analyzes video data to identify emergency situations and recognizes user emotions from audio data.

[1597] Step 4:

[1598] The server's emotion recognition engine analyzes the user's tone of voice, speed, and word choice from the audio data to determine whether the user is anxious, calm, or panicked.

[1599] Step 5:

[1600] Based on the analysis results, the server determines the type and severity of the emergency, selects the most appropriate emergency service contact, and sends this information back to the device.

[1601] Step 6:

[1602] Based on the instructions received, the device automatically calls the selected emergency service contact and plays a server-generated voice message explaining the situation.

[1603] Step 7:

[1604] The device provides first aid instructions to the user. If the server's emotion recognition engine determines that the user's emotional state is impatient or panicked, the device adjusts the content and speed of the instructions to help the user calm down.

[1605] Step 8:

[1606] The server calculates the optimal route to the nearest hospital or fire station based on real-time location information and sends that information to the device, which then displays the calculated route information to the user and begins turn-by-turn navigation.

[1607] Step 9:

[1608] The emotion recognition engine takes into account the user's emotions and provides additional information to emergency contacts, allowing them to make appropriate preparations before emergency services arrive on scene.

[1609] Step 10:

[1610] By following the instructions on the device and receiving first aid and navigation, users can take appropriate action until emergency services arrive, which will enable quick and accurate response and is expected to improve the survival rate.

[1611] Example 2

[1612] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1613] Conventional emergency response systems are required to respond not only based on user reports but also by taking into account the user's emotional state. However, current systems have difficulty accurately recognizing the user's emotions and providing appropriate guidance and emergency measures. Another challenge is providing accurate responses in environments with unstable communication conditions.

[1614] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1615] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call, a guidance means for adjusting first aid guidance according to the user's emotional state, and a means for calculating an optimal route based on the user's real-time location information and providing navigation, thereby enabling a quick and accurate emergency response that also takes the user's emotional state into consideration.

[1616] 1. "Image capture means" refers to a device that activates a camera and captures video or still images when a user presses a trigger.

[1617] 2. "Location information acquisition means" refers to a device that acquires location information of the user's current location using a GPS module or the like.

[1618] 3. "Transmitting means" refers to a device that packetizes the acquired image data, audio data, and location information and transmits them to the server.

[1619] 4. "Analysis means" is a device that analyzes the data transmitted by the transmission means and identifies an emergency situation or the user's emotions.

[1620] 5. "Contact Means" means a device that determines the appropriate emergency contact based on the emergency situation identified by the Analysis Means and automatically initiates a call.

[1621] 6. "Guidance means" is a device that provides the user with first aid guidance and appropriate actions according to their emotional state.

[1622] 7. "Means for providing navigation" means a device that calculates the optimal route based on the user's real-time location information and provides turn-by-turn navigation.

[1623] 8. "Acoustic analysis means" refers to a device that analyzes the surrounding acoustic environment and the user's emotions based on audio data.

[1624] 9. "Language model" refers to a model that is installed to function even in environments with unstable communication and is used to generate emergency and guidance messages.

[1625] MODE FOR CARRYING OUT THE INVENTION

[1626] This invention is an emergency response system that takes into account the user's emotions and can respond quickly and appropriately in an emergency. This system operates via a smartphone app and handles emergencies using an emotion recognition engine and multimodal generation AI.

[1627] Hardware Configuration

[1628] The hardware configuration of this system is as follows:

[1629] 1. Device: A smartphone carried by the user, equipped with the following functions:

[1630] Camera features

[1631] Microphone function

[1632] GPS Modules

[1633] Internet communication function

[1634] 2. Server: A back-end system connected to terminals via the Internet, and includes the following components:

[1635] Multimodal generation AI (audio data analysis, image data analysis)

[1636] Emotion Recognition Engine

[1637] Data Analysis Engine

[1638] Software Configuration

[1639] 1. Terminal application: A smartphone application used by the user that performs the following actions in the event of an emergency.

[1640] Emergency button activates camera and microphone

[1641] Obtaining location information

[1642] Packetizing acquired data (video, audio, location information) and sending it to the server

[1643] 2. Server application: Software that runs on the server side and provides the following functions:

[1644] Analyzing received data

[1645] Recognizing the user's emotional state

[1646] Determining the type and severity of an emergency

[1647] Selection of emergency contacts and automatic call instructions

[1648] Generate and submit a first aid guide

[1649] Real-time location-based navigation

[1650] Example of operation

[1651] To give a concrete example of how this works, consider the following scenario:

[1652] 1. If the emergency button is pressed:

[1653] A user encounters a traffic accident and presses the emergency button on their smartphone.

[1654] The device immediately activates the camera and records video of the accident scene.

[1655] At the same time, the microphone is enabled to record audio and location information is obtained using GPS.

[1656] The terminal transmits this data to the server.

[1657] 2. Data transmission and analysis:

[1658] The server adds the received video, audio, and location information to an analysis queue and activates the multimodal generation AI and emotion recognition engine.

[1659] The server analyzes the user's emotional state (e.g., panic) from the tone of the voice and determines the appropriate emergency response.

[1660] 3. Contacting the appropriate emergency services:

[1661] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and issues an automatic call instruction.

[1662] The device automatically contacts emergency services and plays a server-generated voice message explaining the situation.

[1663] Prompt Sentence Examples

[1664] Examples of prompts include:

[1665] "A user encounters a traffic accident and presses the emergency button. The mobile device collects video, audio, and GPS data and sends it to a server. The server analyzes the urgency of the accident and coordinates the appropriate emergency services. Please explain how the system works."

[1666] This system enables a fast and effective response in emergency situations while taking into consideration the user's feelings, which is expected to improve the user's sense of security and increase the overall survival rate.

[1667] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1668] Step 1:

[1669] The user presses the emergency button.

[1670] Input: The user presses the emergency button.

[1671] Output: The emergency button press signal is transmitted to the terminal.

[1672] The device immediately activates the camera and starts recording video, while simultaneously activating the microphone for audio recording and the GPS module for location information.

[1673] Specific operation: Video recording, audio recording, and location information acquisition of the accident scene are carried out simultaneously.

[1674] Step 2:

[1675] The terminal combines the acquired video file, audio file, and location information into a single data packet and sends it to the server.

[1676] Input: Video data acquired by the camera, audio data recorded by the microphone, and location information acquired by GPS.

[1677] Output: A data packet is generated and sent to the server.

[1678] Specific operation: Various data is packetized and sent to a server via the Internet.

[1679] Step 3:

[1680] The server adds the received data to a queue for analysis and processing, and activates the multimodal generation AI and emotion recognition engine.

[1681] Input: Data packets sent from the terminal.

[1682] Output: The analysis results include the type of emergency and the user's emotional state.

[1683] Specific operation: The received data is added to the analysis queue, and the AI ​​and emotion recognition engine are activated to begin data analysis.

[1684] Step 4:

[1685] The server's emotion recognition engine analyzes the voice data to determine whether the user is expressing emotions such as impatience, fear, or panic.

[1686] Input: Audio data.

[1687] Output: Analysis results showing the user's emotional state.

[1688] Specific operation: Emotion analysis is performed using voice data to determine the user's emotional state.

[1689] Step 5:

[1690] Based on the analysis results, the server determines the type and urgency of the emergency and selects the appropriate emergency service (e.g., police, ambulance, fire department, etc.).

[1691] Input: The type of emergency situation as the analysis result and the user's emotional state.

[1692] Output: Instructions for optimal emergency service selection.

[1693] Specific operation: Based on the analysis results, the level of urgency is determined to be high, and appropriate emergency services are immediately selected.

[1694] Step 6:

[1695] The device receives instructions from the server and automatically calls the selected emergency service contact.

[1696] Input: Instructions from the server.

[1697] Output: Initiation of automatic contact with emergency services.

[1698] Specific operation: The device makes an automatic call and plays a voice message generated by the server to explain the situation.

[1699] Step 7:

[1700] The device provides first aid guidance according to the user's emotional state.

[1701] Input: Analysis results indicating the user's emotional state.

[1702] Output: First aid guide depending on emotional state.

[1703] Specific actions: Provide breathing techniques and first aid instructions in a calm voice.

[1704] Step 8:

[1705] The server calculates the optimal route to the nearest hospital or fire station based on the user's real-time location information.

[1706] Input: The user's real-time location.

[1707] Output: Optimal route information.

[1708] Specific operation: Analyze real-time location information and calculate the optimal route.

[1709] Step 9:

[1710] The device displays optimal route information to the user and provides turn-by-turn navigation.

[1711] Input: Optimal route information.

[1712] Output: Navigation guide.

[1713] Specific behavior: Support the user by providing detailed directions to the nearest hospital.

[1714] Step 10:

[1715] The emotion recognition engine provides additional information to emergency contacts depending on the user's emotions.

[1716] Input: Data indicating the user's emotional state.

[1717] Output: Additional information for emergency contacts.

[1718] Specific behavior: If the user is confused, this information will also be conveyed to emergency responders so that psychological support can be prepared when they arrive at the scene.

[1719] (Application example 2)

[1720] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1721] In the event of an emergency, if a user feels anxious or scared, it can be difficult to respond appropriately. Furthermore, conventional emergency response systems do not take the user's emotional state into account, which can lead to inadequate responses. Furthermore, the system must function even in environments with unstable communications. The present invention aims to solve these problems and provide a system that allows users to respond to emergencies with peace of mind.

[1722] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1723] In this invention, the server includes means for activating an image acquisition means and acquiring location information using a location information acquisition means when triggered by a user, a transmission means for transmitting the acquired image data, audio data, and location information, an analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state, an emotion recognition means for analyzing the user's emotional state identified by the analysis means, a contact means for determining an appropriate emergency contact based on the analysis result of the emotion recognition means and automatically initiating a call, and a guidance means for providing guidance on appropriate first aid according to the user's emotional state. This enables a quick and appropriate emergency response while taking the user's emotional state into consideration.

[1724] "Image capture means" refers to a technical means for capturing an image or video using a device such as a camera when a user activates a trigger.

[1725] "Location information acquisition means" refers to a technical means for determining the current location of a terminal using GPS or other positioning systems.

[1726] The "transmission means" is a means having a function of transmitting the acquired image data, audio data, and location information to a server or other external system.

[1727] "Analysis Means" means the technical means used to identify an emergency situation and determine appropriate emergency contacts based on the Transmitted Data.

[1728] "Emotion recognition means" is a technical means for analyzing the user's emotional state from voice data, etc., and identifying emotions such as impatience or fear.

[1729] A "contact method" is a method that has the ability to automatically initiate a call to an emergency contact based on an identified emergency condition and emotional state.

[1730] The "guidance means" is a technical means for providing a first aid guide in response to an emergency situation faced by a user, taking into consideration the user's emotional state.

[1731] MODE FOR CARRYING OUT THE INVENTION

[1732] To implement this invention, the user's smartphone terminal, the server, and the emergency service work together. The hardware and software used in each step and the processing content thereof will be specifically described below.

[1733] System configuration

[1734] 1. Hardware

[1735] Device (smartphone): Smartphones are equipped with a camera, microphone, and GPS module. These functions are used to acquire images, audio, and location information.

[1736] Server: The server runs a multimodal generation AI and emotion recognition engine, which analyzes the transmitted data.

[1737] 2. Software

[1738] Image capture tool: Software that activates the camera and captures images and videos.

[1739] Audio recording software: Records audio using a microphone and generates audio data.

[1740] Location information acquisition software: Uses the GPS module to determine the device's current location.

[1741] Data transmission software: Collects the acquired data into a single data packet and sends it to a server over the Internet.

[1742] Emotion recognition engine: Analyzes the user's emotional state from voice data and identifies emotions such as impatience and fear. Specifically, it uses an emotion recognition model from the Transformers library.

[1743] Multimodal generative AI: Identifies emergencies based on submitted data and generates appropriate responses. Uses the nlptown / bert-base-multilingual-uncased-sentiment model.

[1744] Contact Method: Automatically call emergency contacts based on the analyzed results.

[1745] Guidance: Provides first aid guidance according to the user's emotional state. Even if communication is unstable, the guidance continues using a pre-installed language model.

[1746] System Operation

[1747] 1. Pressing the emergency button and collecting data

[1748] When a user presses the emergency button, the device's camera automatically activates and begins recording images and videos. At the same time, the microphone activates and records audio. The GPS module also activates and collects location information. This data is packetized within the device and sent to the server.

[1749] 2. Data Analysis and Emotion Recognition

[1750] The server analyzes the received data and identifies the user's emotional state from the user's voice tone and environmental sounds. An emotion recognition engine analyzes the voice data and identifies whether the user is expressing emotions such as anxiety or fear, and determines the appropriate emergency measures.

[1751] Examples:

[1752] When a user encounters a traffic accident and presses the emergency button, the device immediately begins taking video and audio recordings of the accident scene and acquiring location information. The server analyzes this data to understand the circumstances of the accident and the user's emotional state.

[1753] 3. Selecting and contacting emergency services

[1754] Based on the analysis results, the server selects the most appropriate emergency service (e.g., ambulance, police) and automatically calls the contacts. The server generates a voice message explaining the situation.

[1755] 4. First aid information

[1756] The system takes into account the user's emotional state and provides the most appropriate first aid guidance. For example, if a user is in a panic, a slow voice will guide them through CPR and how to treat injuries.

[1757] Examples:

[1758] By inputting the following prompts into the generative AI model, specific emergency response instructions can be generated.

[1759] Example prompt:

[1760] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[1761] As a result, it is possible to take prompt and appropriate emergency measures while taking into consideration the emotional state of the user.

[1762] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1763] Step 1:

[1764] When a user presses the emergency button, the device immediately activates the camera and starts recording images and videos. At the same time, it activates the microphone to record audio and uses the GPS module to obtain location information. These data are packetized and ready to be transmitted. The input is the user pressing the button, and the output is image data, audio data, and location information.

[1765] Step 2:

[1766] The device collects the acquired image data, audio data, and location information into a single data packet and sends it to a server via the Internet. The input is the image data, audio data, and location information, and the output is the data packet sent to the server.

[1767] Step 3:

[1768] The server analyzes the received data packets using a multimodal generation AI and emotion recognition engine. First, it identifies the emergency situation from image data and location information, and then analyzes the audio data to recognize the user's emotional state. The input is the data packets, and the output is the emergency situation identification result and the user's emotional state recognition result.

[1769] Step 4:

[1770] The server analyzes the voice data using an emotion recognition engine to determine whether the user is expressing emotions such as anxiety or fear. For example, it uses an emotion recognition model from the transformers library. The input is the voice data, and the output is the identification of the emotional state.

[1771] Step 5:

[1772] The server determines the appropriate emergency contact based on the analysis results and instructs the device to automatically initiate a call. The server generates the necessary call content using a generative AI model and sends it to the device. The input is the emergency state and emotional state identification results, and the output is the emergency contact and call content.

[1773] Step 6:

[1774] The device receives instructions from the server and automatically calls the appropriate emergency contact. It responds by playing a voice message generated by the server and explaining the situation. The input is the call instruction from the server, and the output is connecting the phone and explaining the situation to the other party.

[1775] Step 7:

[1776] The server uses a generative AI model to create a first aid guide based on the user's emotional state and sends it to the device. The device then provides this guide to the user. Even if communication is unstable, the device continues providing guidance using the installed language model. The input is the user's emotional state and situation data, and the output is the first aid guide content.

[1777] Specific examples

[1778] The following prompts can be fed into a generative AI model to generate specific emergency response instructions:

[1779] Example prompt:

[1780] An emergency has occurred. The user is panicking and at the scene of a traffic accident. Create an appropriate calming voice message and guidelines for first aid at the scene of the accident.

[1781] The above series of steps allows for a quick and appropriate emergency response while taking into account the user's emotional state.

[1782] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1783] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1784] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1785] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1786] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1787] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1788] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1789] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1790] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1791] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1792] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1793] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1794] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1795] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1796] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1797] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1798] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1799] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1800] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1801] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1802] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1803] The following is further disclosed regarding the above embodiment.

[1804] (Claim 1)

[1805] means for activating the image capturing means and acquiring location information using the location information acquiring means when triggered by a user;

[1806] a transmitting means for transmitting the acquired image data, audio data, and location information;

[1807] analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state;

[1808] a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call;

[1809] a guidance means for providing a first aid guide to a user;

[1810] A system including:

[1811] (Claim 2)

[1812] 2. The system according to claim 1, wherein the analyzing means includes acoustic analyzing means for analyzing a surrounding sound environment based on audio data.

[1813] (Claim 3)

[1814] 2. The system of claim 1, wherein the guidance means uses a language model that is installed to function even in an environment where communication is unstable.

[1815] "Example 1"

[1816] (Claim 1)

[1817] means for activating the image capturing means and the audio capturing means and acquiring location information using the location information capturing means when triggered by a user;

[1818] a transmitting means for transmitting the acquired image data, audio data, and location information;

[1819] analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state;

[1820] a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call;

[1821] a guidance means for providing a first aid guide to a user;

[1822] A means for calculating an optimal route based on the user's location information and providing navigation;

[1823] A system including:

[1824] (Claim 2)

[1825] 2. The system according to claim 1, wherein the analysis means uses multimodal generation AI to comprehensively analyze image data, audio data, and location information.

[1826] (Claim 3)

[1827] 2. The system of claim 1, wherein the guidance means uses a language model that is installed to function even in an environment where communication is unstable.

[1828] "Application Example 1"

[1829] (Claim 1)

[1830] means for activating the image capturing means and acquiring location information using the location information acquiring means when triggered by a user;

[1831] a transmitting means for transmitting the acquired image data, audio data, and location information;

[1832] analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state;

[1833] a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call;

[1834] a guidance means for providing a first aid guide to a user;

[1835] a means for calculating an optimal travel route based on a current location of the user and providing navigation;

[1836] A system including:

[1837] (Claim 2)

[1838] 2. The system according to claim 1, wherein the analyzing means includes acoustic analyzing means for analyzing a surrounding sound environment based on audio data.

[1839] (Claim 3)

[1840] 2. The system of claim 1, wherein the guidance means uses a language model that is installed to function even in an environment where communication is unstable.

[1841] "Example 2: Combining Emotion Engines"

[1842] (Claim 1)

[1843] means for activating the image capturing means and acquiring location information using the location information acquiring means when triggered by a user;

[1844] a transmitting means for transmitting the acquired image data, audio data, and location information;

[1845] analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state;

[1846] a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call;

[1847] a guidance means for adjusting the first aid guidance according to the emotional state of the user;

[1848] A means for calculating an optimal route based on real-time location information of a user and providing navigation;

[1849] A system including:

[1850] (Claim 2)

[1851] 2. The system according to claim 1, wherein the analyzing means includes acoustic analyzing means for analyzing the surrounding sound environment and the user's emotions based on the voice data.

[1852] (Claim 3)

[1853] 2. The system of claim 1, wherein the guidance means uses a language model that is installed to function even in an environment where communication is unstable.

[1854] "Application example 2 when combining emotion engines"

[1855] (Claim 1)

[1856] means for activating the image capturing means and acquiring location information using the location information acquiring means when triggered by a user;

[1857] a transmitting means for transmitting the acquired image data, audio data, and location information;

[1858] analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state;

[1859] emotion recognition means for analyzing the emotional state of the user identified by the analysis means;

[1860] a contact means for determining an appropriate emergency contact based on the analysis result of the emotion recognition means and automatically initiating a call;

[1861] a guidance means for providing a guide for appropriate first aid according to the emotional state of the user;

[1862] A system including:

[1863] (Claim 2)

[1864] 2. The system according to claim 1, wherein the analyzing means includes acoustic analyzing means for analyzing a surrounding sound environment based on audio data.

[1865] (Claim 3)

[1866] 2. The system of claim 1, wherein the guidance means uses a language model that is installed to function even in an environment where communication is unstable. [Explanation of symbols]

[1867] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for activating the image capturing means and acquiring location information using the location information acquiring means when triggered by a user; a transmitting means for transmitting the acquired image data, audio data, and location information; analysis means for analyzing the data transmitted by the transmission means and identifying an emergency state; a contact means for determining an appropriate emergency contact based on the emergency state identified by the analysis means and automatically initiating a call; a guidance means for providing a first aid guide to a user; A system including:

2. 2. The system according to claim 1, wherein the analyzing means includes acoustic analyzing means for analyzing the surrounding sound environment based on the audio data.

3. 2. The system of claim 1, wherein the guidance means uses a language model installed to function in an environment where communication is unstable.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A