System
A system using voice and gyro sensor information to detect and respond to emergencies in real time, addressing the lack of timely alerts for the elderly and children by integrating machine learning for anomaly detection and security activation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Existing systems lack the ability to detect anomalies in real time and issue appropriate alerts for emergencies involving the elderly, children, and students, such as falls and kidnappings, necessitating a system that can quickly and accurately use audio and gyro sensor information to protect them.
A system that acquires voice and gyro sensor information, analyzes it for abnormalities using machine learning models, and sends alerts or activates security alarms when necessary, integrating with a server for timely responses.
The system effectively and quickly protects the safety of vulnerable individuals by detecting anomalies and triggering alerts or alarms, ensuring prompt emergency responses.
Smart Images

Figure 2026037451000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, there has been an increase in fraud and crimes targeting the elderly, as well as malicious crimes such as kidnappings targeting children and students. Early detection and rapid response to these crimes is essential, but existing systems lack the ability to detect anomalies in real time or issue appropriate alerts. Therefore, there is a need for a system that uses audio and gyro sensor information to quickly and accurately detect anomalies and issue warnings to users and other relevant parties. [Means for solving the problem]
[0005] The present invention provides a system including means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information and gyro sensor information to a server, means for the server to analyze the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, and means for activating a security buzzer when an abnormality is detected, thereby making it possible to quickly and effectively protect the safety of the elderly and children and students.
[0006] "Audio information" is sound data acquired through the microphone of the terminal.
[0007] "Gyro sensor information" is data on the tilt and acceleration of the device acquired by the device's gyro sensor.
[0008] "Acquisition means" refers to the functions of the equipment or software used to collect audio information and gyro sensor information.
[0009] The "transmission means" refers to the function of the device or software for transmitting the acquired voice information and gyro sensor information to the server.
[0010] "Analysis means" refers to a set of devices and algorithms used by the server to analyze audio information and gyro sensor information to detect abnormalities.
[0011] "Anomaly detection" is the process by which the server analyzes acquired audio and gyro sensor information to check for unexpected behavior or audio patterns.
[0012] "Alert sending" is a means of notifying the user and their relatives of detected abnormalities.
[0013] "Notification sending" is a means of reporting to the police or other relevant authorities when a serious abnormality is detected.
[0014] "Security alarm activation" is a function that activates the security alarm when an abnormality is detected, and issues a warning to those in the vicinity.
[0015] A "machine learning model" is an algorithm trained using past data and used to analyze newly acquired audio information and gyro sensor information.
[0016] "Determining the level of urgency" is the process of evaluating the urgency of a detected abnormality and determining the priority of the response and the destination of the alert. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention is a system that includes a means for acquiring audio information, a means for acquiring gyro sensor information, a means for transmitting the acquired audio information and gyro sensor information to a server, a means for the server to analyze the audio information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected.
[0039] Overview of program processing
[0040] 1. Data Acquisition
[0041] The device (smartphone) acquires voice information using a built-in microphone. This is done using a dedicated application on the device. The voice information is collected at regular intervals and saved as data. The device also acquires information from its built-in gyro sensor in real time to monitor the user's movements and tilt. This gyro sensor information is also saved at regular intervals.
[0042] 2. Sending data to the server
[0043] The voice and gyro sensor information collected by the device is sent to a server in real time or at regular intervals, and the server receives the data and stores it in a database.
[0044] 3. Anomaly detection
[0045] The server analyzes audio and gyro sensor information to detect abnormalities. This analysis is performed using a machine learning model. Audio analysis detects specific keywords (such as screams or "help me"). Gyro sensor analysis also detects sudden movements and unnatural fluctuations (such as falling or sudden changes in acceleration).
[0046] 4. Sending alerts
[0047] If an abnormality is detected, the server evaluates it, and if it is a minor abnormality, it sends an alert to the user and their relatives. This alert is sent via email or app notification. If a serious abnormality is detected, the server also notifies the police, allowing for a prompt emergency response.
[0048] 5. Activating the security alarm
[0049] If an abnormality is detected and is deemed to be particularly serious, the server sends a command to the device to activate the security alarm, thereby notifying people around the user of the abnormality and alerting them to the emergency.
[0050] Specific examples
[0051] Below is a concrete example of how the system of the present invention actually works.
[0052] Consider the case where an elderly person suddenly collapses. At this time, the gyro sensor on the device detects the sudden movement and sends the gyro sensor information to the server as a falling motion. The server analyzes the received gyro information and detects an abnormality such as a fall. If the server determines that this abnormality is serious, it notifies the family and also reports the incident to the police. The server also sends instructions to the device to activate the security alarm. This alerts people in the vicinity to the abnormality, enabling a swift response.
[0053] The system also simulates a child being abducted. In this case, the gyro sensor detects any sudden movements or changes in behavior, and the child's cries for help are captured as audio information. This information is sent to the server, which analyzes it to detect any abnormalities. The server determines this to be a very serious anomaly and immediately notifies the police. At the same time, a notification is sent to the parents, and a security alarm is activated. This is expected to allow for early response.
[0054] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0058] Step 2:
[0059] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0060] Step 3:
[0061] The device sends the collected voice and gyro sensor information to a server. In particular, this data is sent to the server in real time or at specified intervals via internet communication. The API used for transmission is encrypted as a security measure.
[0062] Step 4:
[0063] The server stores the received audio and gyro sensor information in a database, which allows analysis to detect anomalies by comparing it with past data.
[0064] Step 5:
[0065] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[0066] Step 6:
[0067] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[0068] Step 7:
[0069] The server determines whether an abnormality is detected based on the results of voice analysis and gyro sensor analysis, and if so, evaluates whether the abnormality is minor or serious.
[0070] Step 8:
[0071] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[0072] Step 9:
[0073] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[0074] Step 10:
[0075] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[0076] Step 11:
[0077] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[0078] Example 1
[0079] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0080] In modern society, when vulnerable people such as the elderly and children face an emergency, there is a need to respond quickly and appropriately, but there are still few systems that can achieve this, and existing technologies are insufficient. In particular, there is a lack of means to detect serious abnormalities such as falls and kidnappings early and to take prompt and appropriate action.
[0081] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0082] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information and gyro sensor information to the server, means for analyzing the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, means for activating a security buzzer when an abnormality is detected, means for sending the voice information and gyro sensor information acquired by the terminal to the server via Wi-Fi or a mobile data network, and means for collecting and storing various sensor information at time intervals. This makes it possible to quickly collect and analyze voice information and gyro sensor information and take appropriate action when elderly people, children, or others face an emergency.
[0083] "Means for acquiring audio information" refers to the means by which a terminal uses a built-in microphone or other audio recording device to collect surrounding audio as digital data.
[0084] "Means for acquiring gyro sensor information" refers to the means by which a device uses its built-in gyro sensor to collect data on the user's movements and tilt in real time.
[0085] "Means for transmitting acquired voice information and gyro sensor information to a server" refers to means for transmitting voice information and gyro sensor information collected by the terminal to a server via a communication means such as Wi-Fi or a mobile data network.
[0086] "Means for the server to analyze audio information and gyro sensor information to detect anomalies" refers to means for the server to use machine learning models or other analytical technologies to analyze received audio information and gyro sensor information and detect anomalies.
[0087] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to the means for informing the user and their relatives of an abnormality by sending an email or app notification when the server detects an abnormality.
[0088] The "means for sending a notification to the police when a serious abnormality is detected" is a means for the server to quickly notify the police when it detects a serious abnormality.
[0089] The "means for activating the security buzzer when an abnormality is detected" is a means for the server to send an instruction to the terminal to activate the security buzzer of the terminal when an abnormality is detected.
[0090] "Means for transmitting voice information and gyro sensor information acquired by a terminal to a server via Wi-Fi or a mobile data network" refers to means for transmitting voice information and gyro sensor information collected by a terminal to a server using wireless communication technology.
[0091] "Means for collecting and storing various sensor information at regular intervals" refers to means by which the terminal automatically collects and stores audio information and gyro sensor information at regular intervals.
[0092] This invention is a system for responding quickly and appropriately to emergencies faced by vulnerable people such as the elderly and children. This system acquires voice and gyro sensor information, transmits it to a server for analysis, and activates an alert or a security buzzer when an abnormality is detected.
[0093] Data Acquisition
[0094] The device (smartphone) acquires audio information using a built-in microphone. This is done using a dedicated application on the smartphone. The audio information is collected at regular intervals (for example, every 5 seconds) and saved as data. The device also uses a built-in gyro sensor to monitor the user's movements and tilt in real time. The gyro sensor information is also saved at regular intervals. In this way, the device can acquire audio data and gyro sensor data simultaneously.
[0095] Sending data to the server
[0096] The voice and gyro sensor information collected by the device is sent to a server via Wi-Fi or mobile data network in real time or at regular intervals. The server receives the sent data and stores it in a database. During this process, communication between the device and the server is encrypted to ensure security.
[0097] Anomaly detection
[0098] The server analyzes the received voice information and gyro sensor information. This analysis uses generative AI models and machine learning models. Specifically, natural language processing technology is used to analyze the voice information to detect specific keywords (e.g., "help" or "danger"). Additionally, analysis of the gyro sensor information detects sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration). If an abnormality is detected as a result of the analysis, the system proceeds to the next stage.
[0099] Sending alerts
[0100] If the server detects an abnormality, it evaluates the urgency of that abnormality. If a minor abnormality is detected (for example, if only audio containing specific keywords is detected), the server sends an alert to the user and their relatives. This alert is sent via email or a dedicated application. If a serious abnormality is detected (for example, if audio information and abnormal behavior are detected simultaneously), the server also notifies the police. This allows for a swift emergency response.
[0101] Activating the security alarm
[0102] If the server detects a serious abnormality, it will send a command to the device to activate the security alarm. Activating the security alarm will alert people around the user to the abnormality and make it clear that it is an emergency. This function will enable people around the device to take early action.
[0103] Specific examples
[0104] Falls in the elderly
[0105] If an elderly person suddenly falls, the device's gyro sensor detects the sudden movement and sends that information to the server. The server analyzes the received gyro sensor information and detects the abnormality as a fall. If the server determines that the abnormality is serious, it notifies the family and also reports the incident to the police. The server then sends an instruction to activate the device's security alarm, which then emits a high-volume alarm to alert those around it to the abnormality.
[0106] Child abduction
[0107] In the event of a child abduction, the device's gyro sensor will detect any sudden movements or changes in behavior, and the child's cries for help will be picked up as audio information. This information is sent to a server, which analyzes it to detect any abnormalities. The server will determine this to be a very serious anomaly and immediately notify the police. At the same time, a notification will be sent to the parents, and a security alarm will be activated. This is expected to allow for early response.
[0108] Prompt Sentence Examples
[0109] "Please explain a system that can acquire voice and gyro sensor information and detect abnormalities when an elderly person suddenly collapses."
[0110] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0112] Step 1: Get the data
[0113] The device uses a built-in microphone to acquire audio information. This is done using a dedicated application, and audio is collected at regular intervals (e.g., every 5 seconds). The input is ambient audio, and digitized audio data is obtained as output. The device also uses a built-in gyro sensor to monitor the user's movement and tilt. This gyro sensor information is also collected at regular intervals, and acceleration data for the x-, y-, and z-axes is obtained as output.
[0114] Step 2: Send data to the server
[0115] The device sends the collected voice information and gyro sensor information to the server via Wi-Fi or mobile data network. The input is the voice data and gyro sensor data collected by the device, and this data is sent to the server as packets at regular time intervals (e.g., 5 seconds). The output is the data packets sent to the server.
[0116] Step 3: Detect anomalies
[0117] The server analyzes the received voice information and gyro sensor information. Generative AI models and machine learning models are used for this analysis. Based on the voice data input, natural language processing techniques are used to detect specific keywords (e.g., "help," "danger"), and based on the gyro sensor data input, sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration) are detected. The analysis results are obtained as output, and a flag is raised if an abnormality is detected.
[0118] Step 4: Sending an alert
[0119] If an anomaly is detected, the server evaluates the anomaly and sends an alert to the user and their relatives. The input is the analysis result, and if the anomaly flag is true, an action is triggered. If the anomaly is minor, the output is an email or app notification sent to the user and their relatives. If a serious anomaly is detected, a notification is also sent to the police.
[0120] Step 5: Activate the security alarm
[0121] If a serious abnormality is detected, the server sends an instruction to the terminal to activate the security buzzer. The input is the abnormality evaluation result, and if the serious abnormality flag is true, an instruction to activate the security buzzer is sent as output to the terminal. Upon receiving this instruction, the terminal activates the security buzzer and emits a high-volume alarm to alert those in the vicinity of the abnormality.
[0122] (Application example 1)
[0123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0124] In recent years, ensuring the safety of vulnerable people, including the elderly and children, has become an important issue. However, current security systems often lack the ability to recognize abnormalities in real time or respond quickly. Furthermore, few systems that detect abnormalities using both audio and motion information efficiently integrate these data, resulting in problems such as false detection and delayed response. The present invention aims to solve these problems by providing a system that can quickly and accurately detect and respond to abnormalities, especially when the elderly or children are in danger.
[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0126] In this invention, the server includes means for analyzing voice information and gyro sensor information to detect abnormalities, means for analyzing specific keywords from the voice information and detecting sudden changes in movement from the gyro sensor information, and means for sending an alert and activating a security buzzer if necessary if an abnormality is detected based on the analysis results. This makes it possible to comprehensively monitor the user's voice and movement, respond quickly in the event of an abnormality, and notify the user and their relatives or those around them of an emergency by activating the security buzzer.
[0127] "Audio information" is data that describes the user's environmental sounds and speech content.
[0128] A "gyro sensor" is a device used to detect the user's movement and tilt.
[0129] The "server" is a computer system that analyzes collected audio information and gyro sensor information, and detects and notifies users of abnormalities.
[0130] An "alert" is a notification sent to the user, their relatives, and, if necessary, the police or other relevant parties when an abnormality is detected.
[0131] A "security buzzer" is a device that emits a sound when an abnormality is detected to alert those around to danger.
[0132] "Analysis" is a process for detecting the presence or absence of abnormalities based on collected audio information and gyro sensor information.
[0133] The "specific keywords" are words and phrases that mainly indicate an emergency in the audio information that indicates an abnormality.
[0134] "Sudden changes in movement" refers to sudden body movements that deviate from normal movement.
[0135] The present invention provides a system that uses voice information and gyro sensor information to monitor the safety of a user and responds quickly when an abnormality is detected. Specific embodiments for carrying out the present invention will be described below.
[0136] System Configuration
[0137] Hardware:
[0138] Smartphone: Equipped with a built-in microphone and gyro sensor, it acquires voice and gyro sensor information.
[0139] Server: Analyzes audio and gyro sensor information to detect abnormalities, send notifications, activate the security alarm, etc.
[0140] software:
[0141] Dedicated application: Equipped with functions to collect voice information and gyro sensor information and send it to a server.
[0142] Cloud server: Data analysis and management is performed using AWS (registered trademark), GCP, Azure (registered trademark), etc.
[0143] Machine learning model: Using Python or similar software, an algorithm is implemented to analyze audio information and gyro sensor information and detect anomalies.
[0144] Speech recognition technology: Using technologies such as IBM Watson® and Google® Speech-to-Text, voice data is converted into text and specific keywords are detected.
[0145] Program Processing Overview
[0146] Acquiring audio information
[0147] The user's smartphone uses a built-in microphone to capture voice information, which is then sent to a server at regular intervals via a dedicated application.
[0148] Get gyro sensor information
[0149] The smartphone's built-in gyro sensor is used to monitor the user's movements and tilt in real time, and the application periodically sends the data to a server.
[0150] Data analysis on the server
[0151] The server analyzes the received voice and gyro sensor information. The voice information is converted into text using speech recognition technology to detect whether it contains specific keywords (such as "help" or screams). The gyro sensor information is analyzed using a machine learning model to detect sudden changes in movement (e.g., falling).
[0152] Sending alerts
[0153] If an abnormality is detected, the server evaluates its severity, and if it is a minor abnormality, it sends an app notification to the user and their relatives. If it is a serious abnormality, it also notifies the police.
[0154] Activating the security alarm
[0155] If a serious abnormality is detected, a security alarm will be activated on the smartphone, alerting people nearby.
[0156] Specific examples
[0157] Falls in the elderly
[0158] If an elderly person suddenly falls, the smartphone's gyro sensor will detect the sudden movement and send the information to the server. The server will then detect the fall, notify the family, and notify the police. It will also activate a security alarm to alert people nearby of the emergency.
[0159] Child abduction prevention
[0160] In the case of a child being kidnapped, the smartphone's gyro sensor will detect any sudden changes in movement and capture any voices crying for help. This information will be sent to a server, which will then analyze it and determine that the child has been kidnapped. The server will then immediately notify the police, and a notification will be sent to the parents, who will then activate the security alarm. This is expected to allow for early response.
[0161] Prompt Sentence Examples
[0162] I'd like to see some code examples for analyzing data from an app to detect anomalies in user voice and behavior. The voice data needs to detect specific keywords, such as "help" or "scream." Also, the gyro sensor information needs to detect sudden changes in behavior. I'll be using Python.
[0163] As described above, the system of the present invention can quickly and effectively protect the safety of elderly people and children by using voice information and gyro sensor information.
[0164] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0165] Step 1:
[0166] Acquiring audio information
[0167] The user acquires audio information using a smartphone. The smartphone's built-in microphone collects the user's ambient sounds and speech, and saves the audio data at regular intervals through a dedicated application. The input is the ambient sounds and the user's speech, and the output is the saved audio data.
[0168] Step 2:
[0169] Get gyro sensor information
[0170] The device uses a built-in gyro sensor to monitor the user's movements in real time. The gyro sensor captures acceleration information along the x-, y-, and z-axes, and periodically saves this as data through a dedicated application. The input is the user's movements, and the output is the captured gyro sensor information.
[0171] Step 3:
[0172] Voice information and gyro sensor information sent to the server
[0173] The smartphone transmits the stored voice and gyro sensor information to the server at regular intervals using a secure communication protocol (such as HTTPS). The input is the voice data and gyro sensor information, and the output is the data transmitted to the server.
[0174] Step 4:
[0175] Saving data to the server
[0176] The server stores the received audio information and gyro sensor information in a database. The input is the transmitted data, and the output is the data stored in the database.
[0177] Step 5:
[0178] Analysis of audio information
[0179] The server uses speech recognition technology to convert the stored voice information into text, and analyzes this text data to see if it contains specific keywords (such as "help" or a cry). The input is the stored voice information, and the output is the text data and the analysis results.
[0180] Step 6:
[0181] Analysis of gyro sensor information
[0182] The server analyzes the gyro sensor information using a machine learning model, detects sudden changes in movement (e.g., falling or sudden acceleration), and evaluates the user's situation. The input is the stored gyro sensor information, and the output is the analysis result.
[0183] Step 7:
[0184] Anomaly detection
[0185] The server detects abnormalities based on the results of analyzing voice information and gyro sensor information. It evaluates the degree of abnormality based on the detection results and decides how to respond. The input is the results of voice analysis and gyro analysis, and the output is the presence or absence of an abnormality and its degree.
[0186] Step 8:
[0187] Sending alerts
[0188] The server sends an alert to the user and their relatives depending on the detected anomaly. If the anomaly is minor, it sends an app notification, and if it is severe, it also notifies the police. The input is the degree of anomaly, and the output is the alert sent.
[0189] Step 9:
[0190] Activating the security alarm
[0191] If the server determines that a serious abnormality exists, it sends a command to the smartphone to activate the security buzzer. This notifies people around the smartphone of the emergency. The input is the degree of abnormality, and the output is the activation of the security buzzer.
[0192] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0193] The present invention is a system that includes a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected, and further includes an emotion engine that recognizes the user's emotions.
[0194] Overview of program processing
[0195] 1. Data Acquisition
[0196] The device (smartphone) acquires voice information using a built-in microphone. The voice information is collected at regular intervals and temporarily stored within the device. Motion data is also acquired using a gyro sensor and is also stored within the device.
[0197] 2. Emotion recognition
[0198] The device transmits the acquired voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with other sensor data.
[0199] 3. Sending data to the server
[0200] The device sends collected voice information, gyro sensor information, and emotion information to a server, which sends this data in real time or at specified intervals.
[0201] 4. Anomaly Detection
[0202] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models to detect whether specific keywords (e.g., "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data. It also evaluates whether the user is under high stress based on the emotion information provided by the emotion engine.
[0203] 5. Sending alerts
[0204] If an anomaly is detected, the server will trigger an appropriate alert depending on the severity: if it is a minor anomaly, an alert will be sent to the user and their relatives; if it is a serious anomaly, a notification will also be sent to the police.
[0205] 6. Activating the security alarm
[0206] If a serious abnormality is detected, the server sends a command to the device to activate the burglar alarm, which then automatically sounds the alarm to alert people in the vicinity.
[0207] Specific examples
[0208] For example, let's simulate an elderly person suddenly collapsing at home. The device's gyro sensor detects the sudden movement and sends gyro data to the server as a falling motion. At the same time, a voice calling for help is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server, and the analysis results indicate a serious abnormality. The server notifies the family and also alerts the police. It also sends instructions to the device to activate the security alarm. This allows people in the vicinity to be quickly notified of the abnormality.
[0209] The system also simulates a situation where a child is about to be kidnapped in a park. In this case, the device detects sudden movements and the voice calling for help, while the emotion engine detects strong fear. This information is sent to the server, which determines it as a serious abnormality. The server then immediately notifies the police and parents and instructs the device to activate the security alarm, enabling early action to be taken and helping to ensure the child's safety.
[0210] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children by integrating voice information, gyro sensor information, and emotional information.
[0211] The processing flow will be explained below.
[0212] Step 1:
[0213] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0214] Step 2:
[0215] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0216] Step 3:
[0217] The device sends the collected voice information to the emotion engine, which analyzes the voice data in real time and recognizes the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is temporarily stored in the device along with other sensor data.
[0218] Step 4:
[0219] The device transmits voice information, gyro sensor information, and emotion information to the server. This data is sent to the server in real time or at specified intervals. The API used for transmission is encrypted as a security measure.
[0220] Step 5:
[0221] The server stores the received voice, gyro sensor, and emotion information in a database, which can then be compared with past data for analysis to detect anomalies.
[0222] Step 6:
[0223] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[0224] Step 7:
[0225] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[0226] Step 8:
[0227] The server analyzes the emotional information and evaluates whether the user is under high stress or fear based on the data provided by the emotion engine.
[0228] Step 9:
[0229] The server uses the results of voice, gyro sensor, and emotion analysis to determine whether an anomaly has been detected, and if so, evaluates whether the anomaly is minor or major.
[0230] Step 10:
[0231] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[0232] Step 11:
[0233] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[0234] Step 12:
[0235] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[0236] Step 13:
[0237] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[0238] Example 2
[0239] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0240] Ensuring the safety of elderly people and children in modern society is a critical issue, requiring rapid and accurate responses, especially in emergencies. Conventional anomaly detection systems respond based solely on voice and motion information, and do not take into account changes in emotions. Such systems have difficulty determining the level of urgency, and there is a risk of delays in issuing appropriate alerts or notifications. Therefore, in order to more reliably protect the personal safety of elderly people and children, there is a need to develop a system that can also comprehensively analyze emotional information.
[0241] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0242] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information to an emotion engine to analyze the user's emotion, means for transmitting the analyzed emotion information to the server, means for transmitting the voice information and gyro sensor information to the server, means for the server to analyze the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, and means for activating a security buzzer when an abnormality is detected. This makes it possible to comprehensively analyze the user's emotion information, appropriately determine which abnormalities are of high urgency, and respond quickly.
[0243] The "means for acquiring voice information" is a means for collecting voice data using a built-in microphone or other voice input device.
[0244] "Means for acquiring gyro sensor information" refers to a means for using a gyro sensor to measure fluctuations in the user's movements and posture and collect that data.
[0245] The "means for transmitting acquired voice information and gyro sensor information to a server" refers to a means having the function of transmitting collected voice data and gyro data to a server via wireless communication or cable connection.
[0246] "Means for the server to analyze audio information and gyro sensor information to detect abnormalities" refers to means for analyzing audio data and gyro data within the server and executing algorithms or programs that detect abnormalities based on the results.
[0247] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to a means that has the function of sending a notification to the user and their pre-designated relatives when an abnormality is detected.
[0248] "Means for sending a notification to the police when a serious abnormality is detected" refers to means that has the function of sending an emergency notification to public institutions such as the police when a highly urgent abnormality is detected.
[0249] "Means for activating the security alarm when an abnormality is detected" refers to means having the function of sounding the security alarm of the terminal to notify those around when the system detects an abnormality.
[0250] An "emotion engine that recognizes user emotions" is an algorithm or program that analyzes acquired voice data to identify and identify the user's emotional state.
[0251] The "means for transmitting acquired voice information to an emotion engine and analyzing the user's emotions" refers to a means having a function for transmitting collected voice data to an emotion engine and analyzing the emotional state.
[0252] The "means for transmitting analyzed emotion information to a server" is a means having a function for transmitting emotion data analyzed by an emotion engine to a server.
[0253] The system of the present invention collects user voice and gyro sensor information, sends it to a server for analysis, detects abnormalities, and issues appropriate alerts and notifications. It also has the ability to collect and analyze user emotional information to more accurately determine the urgency of the abnormality.
[0254] Hardware and software used
[0255] This system uses the following hardware and software:
[0256] Hardware:
[0257] 1. Device (smartphone): Collects voice information using the built-in microphone and acquires movement data using the built-in gyro sensor.
[0258] 2. Server: Analyzes data, detects anomalies, and manages alerts and notifications.
[0259] software:
[0260] 1. Emotion engine: Software for emotion recognition (e.g., IBM Watson or Google Cloud Speech-to-Text).
[0261] 2. Machine learning model: Anomaly detection model trained using TENSORFLOW® and PyTorch.
[0262] 3. Data transmission and management application: An application that manages data transmission from the terminal to the server and instructions from the server.
[0263] Data Acquisition and Processing
[0264] The device uses a built-in microphone to collect audio information at regular intervals (e.g., every 5 seconds). The collected audio data is temporarily stored in the device. Similarly, the built-in gyro sensor is used to obtain motion data, which is also stored in the device.
[0265] The device then sends the collected voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with the collected gyro data.
[0266] The server uses voice analysis algorithms and machine learning models to analyze the received voice information, gyro data, and emotional information. If an abnormality is detected, the server triggers an appropriate alert depending on the urgency. If the abnormality is minor, the server notifies the user and their relatives, and if it is serious, it notifies the police. If a serious abnormality is detected, the server sends an instruction to the device to sound the security alarm.
[0267] Specific examples
[0268] Example 1: An elderly person suddenly collapses at home
[0269] The device's gyro sensor detects violent movements and sends gyro data to the server as a falling motion. At the same time, voice information calling for "help" is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server in real time. The server analyzes this information using a voice analysis algorithm and machine learning model and determines that it is a serious abnormality. The server then sends a notification to family members and alerts the police. It can also send instructions to the device to activate a security alarm, quickly alerting people in the vicinity of an abnormality.
[0270] Example 2: A child is nearly kidnapped in a park
[0271] The device detects sudden movements and cries for help, while the emotion engine detects strong fear. This information is sent to the server in real time. The server immediately detects any abnormalities and notifies the police and parents as serious problems. The server also sends instructions to the device to sound a security alarm, alerting people in the vicinity. This encourages rapid response and helps ensure the safety of children.
[0272] Prompt Sentence Examples
[0273] Example prompt: "Write an outline of an emergency alert system for people with disabilities. The system has voice recognition, emotion detection, and anomaly detection capabilities, and the ability to notify users of detected anomalies."
[0274] Sample prompt: "Describe an anomaly detection system for when an elderly person suddenly collapses. The system uses voice, gyro sensor, and emotion data to detect the anomaly."
[0275] In this way, the system of the present invention makes it possible to quickly and effectively protect the safety of elderly people and children by integrating voice information, gyro sensor information, and emotion information.
[0276] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0277] Step 1:
[0278] Data Acquisition
[0279] The device (smartphone) acquires audio information using a built-in microphone. The microphone continuously captures audio and collects data at regular intervals (for example, every 5 seconds). The input is the user's voice, and the output is audio data stored in the device. Similarly, the device's built-in gyro sensor captures the user's movement data in real time. Gyro data is also collected at regular intervals and stored in the device. The input is the user's movement, and the output is gyro sensor data.
[0280] Step 2:
[0281] emotion recognition
[0282] The device sends the acquired voice information in real time to an emotion engine (for example, IBM Watson or Google Cloud Speech-to-Text) to recognize the user's emotions. The input is the voice information stored on the device, and the output is the emotional state (anger, anxiety, sadness, etc.) obtained from the emotion engine. For example, if the user is angry, the emotion engine will recognize "anger" based on the tone and content of the voice. This emotion information is temporarily stored on the device and sent to the server along with gyro data.
[0283] Step 3:
[0284] Sending data to the server
[0285] The device sends voice information, gyro sensor information, and emotion information to the server. The input is the voice data, gyro data, and emotion information stored on the device, and the output is each data sent to the server. This data transmission is done in real time, but depending on the communication environment, it may also be sent at specified time intervals (for example, every minute).
[0286] Step 4:
[0287] Anomaly detection
[0288] The server analyzes the received voice information, gyro sensor information, and emotional information. The input is the voice data, gyro data, and emotional information sent to the server, and the output is the anomaly detection results. The server is equipped with a machine learning model (for example, a model trained using TensorFlow or PyTorch) and uses this to detect anomalies. Specifically, the server uses a voice analysis algorithm to detect whether the voice data contains specific keywords (for example, "help" or screams). It also analyzes patterns in the gyro data to determine whether the user has suddenly fallen. Furthermore, it evaluates whether the user is experiencing strong stress or fear based on the emotional information provided by the emotion engine.
[0289] Step 5:
[0290] Sending alerts
[0291] If an anomaly is detected, the server triggers an alert depending on its urgency. The input is the anomaly detection result, and the output is the alert to be sent. If the anomaly is minor (for example, a small noise or a slight fall), the server will send an SMS or app notification to the user and their relatives. If the anomaly is serious (for example, if a cry for help or strong emotional expression is detected), the server will immediately notify the police. The notification also includes location information.
[0292] Step 6:
[0293] Activating the security alarm
[0294] When a serious abnormality is detected, the server sends a command to the device to activate the security buzzer. The input is the serious abnormality detection result, and the output is the sound of the security buzzer. The device receives this command and sounds the security buzzer using its built-in speaker. This alerts people in the vicinity to the abnormality and encourages them to take early action.
[0295] (Application example 2)
[0296] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0297] Conventional security systems rely solely on audio and motion information to detect abnormalities, resulting in low accuracy in detecting abnormalities and often making it difficult to determine whether an abnormality has occurred because they do not take into account the emotional state of the user. Furthermore, even if an abnormality is detected, prompt action may not be possible, posing a challenge to ensuring user safety.
[0298] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0299] The smart glasses include a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, a means for activating a security buzzer when an abnormality is detected, a means including an emotion engine that recognizes the user's emotions, a means for transmitting emotion information recognized by the emotion engine to a server, and a means for supporting the safety of users using the smart glasses. This improves the accuracy of abnormality detection and enables prompt and appropriate responses.
[0300] A "means for acquiring voice information" is a device or mechanism that detects a user's voice and records and collects it as digital data.
[0301] "Means for acquiring gyro sensor information" refers to devices or mechanisms that measure changes in a user's movements and posture, and record and collect them as digital data.
[0302] The "means for transmitting to a server" refers to a device or method for transmitting the acquired voice information and gyro sensor information to a server via a communication means such as wireless communication or the Internet.
[0303] "Means for analyzing and detecting anomalies" refers to a processing system that analyzes the data received by the server using algorithms, machine learning models, etc., to identify anomalies.
[0304] An "alert sending means" is a device or method for sending a warning or caution notification to a designated recipient when an abnormality is detected.
[0305] "Means for sending notifications" refers to devices or methods for promptly contacting the user's relatives, police, or other relevant authorities in the event of a serious abnormality.
[0306] "Means for activating a security alarm" refers to the mechanism or control method for activating a physical audio alarm device (security alarm) when an abnormality is detected.
[0307] An "emotion engine" is an algorithm or software that analyzes acquired voice information and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.).
[0308] "Smart glasses" are high-performance glasses that can acquire audio information and gyro sensor information, and can also recognize the user's emotions using an emotion engine, displaying visual information and communicating.
[0309] In this invention, we will build a system in which smart glasses and other applicable devices acquire voice information and gyro sensor information and send it to a server. This allows for quick and appropriate response when an abnormality occurs. Specific implementation methods are shown below.
[0310] First, we use smart glasses as the hardware. The smart glasses have built-in microphones and gyro sensors, which enable them to collect voice and motion information. They also have wireless communication capabilities such as Wi-Fi and Bluetooth, which allow them to send collected data to a server. Second, we use Google Cloud's Speech-to-Text API and Emotion API as emotion engines for emotion recognition.
[0311] The data received by the server is:
[0312] 1. Audio information: The user's voice is picked up by a microphone and recorded as digital data.
[0313] 2. Gyro sensor information: Detects user movements and collects their movement data.
[0314] 3. Emotion information: Emotion data analyzed by the emotion engine.
[0315] The server analyzes the received data and detects anomalies using machine learning models (for example, TensorFlow or PyTorch). Based on the analysis results, it determines the urgency of the anomaly and sends an appropriate alert. Specifically, it notifies the user and their relatives using Twilio's SMS API or Push Notification API.
[0316] If the abnormality is determined to be serious, the server sends a notification to the police and instructs the device to activate a security alarm, which activates a physical audio alarm and alerts people in the vicinity to the abnormality.
[0317] As a concrete example, consider the case where an elderly person suddenly falls while out. In this case, the gyro sensor in the smart glasses detects the fall, and the microphone picks up the cry of "Ouch!" The emotion engine distinguishes between strong pain and fear, and sends this data to the server. The server determines this to be a serious abnormality, sends an emergency call to relatives, and activates the security alarm.
[0318] Prompt Sentence Examples
[0319] "Please build a system that collects voice and gyro sensor information in real time to detect when the user falls or has a sudden change in behavior, and sends the data along with the user's emotional state to a server for analysis and alerts when an abnormality occurs."
[0320] In this way, the invention provides a system that uses smart glasses to collect voice information and gyro sensor information, and also integrates and analyzes emotional information, thereby improving the accuracy of anomaly detection and enabling prompt and appropriate responses.
[0321] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0322] Step 1: Acquire audio and gyro sensor information
[0323] The device (smart glasses) collects the user's voice information using a built-in microphone. At the same time, it acquires the user's movement data using a gyro sensor. The acquired voice information and gyro sensor information are temporarily stored in the device. This input data is the raw data for monitoring the user's voice and movement in real time.
[0324] Step 2: Emotion Recognition
[0325] The device sends the collected voice information to an emotion engine in real time to analyze the user's emotions. The emotion engine (for example, Google Cloud's Speech-to-Text API or Emotion API) analyzes the voice data and identifies the emotional state (anger, anxiety, sadness, etc.). It receives voice information as input data and outputs emotional information. This emotional information is also temporarily stored on the device.
[0326] Step 3: Send data
[0327] The device transmits the acquired voice information, gyro sensor information, and emotion information to the server via Wi-Fi or Bluetooth. This data is sent to the server as input data for analysis.
[0328] Step 4: Anomaly detection
[0329] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models (TensorFlow and PyTorch) to detect whether specific keywords (such as "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the data patterns from the gyro sensor. It also evaluates the user's level of stress based on the emotion information provided by the emotion engine. The analysis results indicate whether there is an abnormality and its urgency.
[0330] Step 5: Sending an alert
[0331] If an abnormality is detected, an alert is sent according to the urgency. If the abnormality is minor, the server sends an alert to the user and their relatives. SMS API and Push Notification API are used for sending the alert. If a serious abnormality occurs, the server also sends a notification to the police. The input data is the analysis result, and the output data is the alert notification.
[0332] Step 6: Activate the security alarm
[0333] If a serious abnormality is detected, the server sends an instruction to the device to activate the security alarm. This activates the security alarm in the smart glasses and alerts people nearby. The input data is the instruction from the server, and the output data is the activation of the security alarm.
[0334] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0335] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0336] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0337] [Second embodiment]
[0338] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0339] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0340] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0341] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0342] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0343] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0344] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0345] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0346] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0347] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0348] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0349] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0350] The present invention is a system that includes a means for acquiring audio information, a means for acquiring gyro sensor information, a means for transmitting the acquired audio information and gyro sensor information to a server, a means for the server to analyze the audio information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected.
[0351] Overview of program processing
[0352] 1. Data Acquisition
[0353] The device (smartphone) acquires voice information using a built-in microphone. This is done using a dedicated application on the device. The voice information is collected at regular intervals and saved as data. The device also acquires information from its built-in gyro sensor in real time to monitor the user's movements and tilt. This gyro sensor information is also saved at regular intervals.
[0354] 2. Sending data to the server
[0355] The voice and gyro sensor information collected by the device is sent to a server in real time or at regular intervals, and the server receives the data and stores it in a database.
[0356] 3. Anomaly detection
[0357] The server analyzes audio and gyro sensor information to detect abnormalities. This analysis is performed using a machine learning model. Audio analysis detects specific keywords (such as screams or "help me"). Gyro sensor analysis also detects sudden movements and unnatural fluctuations (such as falling or sudden changes in acceleration).
[0358] 4. Sending alerts
[0359] If an abnormality is detected, the server evaluates it, and if it is a minor abnormality, it sends an alert to the user and their relatives. This alert is sent via email or app notification. If a serious abnormality is detected, the server also notifies the police, allowing for a prompt emergency response.
[0360] 5. Activating the security alarm
[0361] If an abnormality is detected and is deemed to be particularly serious, the server sends a command to the device to activate the security alarm, thereby notifying people around the user of the abnormality and alerting them to the emergency.
[0362] Specific examples
[0363] Below is a concrete example of how the system of the present invention actually works.
[0364] Consider the case where an elderly person suddenly collapses. At this time, the gyro sensor on the device detects the sudden movement and sends the gyro sensor information to the server as a falling motion. The server analyzes the received gyro information and detects an abnormality such as a fall. If the server determines that this abnormality is serious, it notifies the family and also reports the incident to the police. The server also sends instructions to the device to activate the security alarm. This alerts people in the vicinity to the abnormality, enabling a swift response.
[0365] The system also simulates a child being abducted. In this case, the gyro sensor detects any sudden movements or changes in behavior, and the child's cries for help are captured as audio information. This information is sent to the server, which analyzes it to detect any abnormalities. The server determines this to be a very serious anomaly and immediately notifies the police. At the same time, a notification is sent to the parents, and a security alarm is activated. This is expected to allow for early response.
[0366] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0367] The processing flow will be explained below.
[0368] Step 1:
[0369] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0370] Step 2:
[0371] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0372] Step 3:
[0373] The device sends the collected voice and gyro sensor information to a server. In particular, this data is sent to the server in real time or at specified intervals via internet communication. The API used for transmission is encrypted as a security measure.
[0374] Step 4:
[0375] The server stores the received audio and gyro sensor information in a database, which allows analysis to detect anomalies by comparing it with past data.
[0376] Step 5:
[0377] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[0378] Step 6:
[0379] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[0380] Step 7:
[0381] The server determines whether an abnormality is detected based on the results of voice analysis and gyro sensor analysis, and if so, evaluates whether the abnormality is minor or serious.
[0382] Step 8:
[0383] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[0384] Step 9:
[0385] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[0386] Step 10:
[0387] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[0388] Step 11:
[0389] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[0390] Example 1
[0391] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0392] In modern society, when vulnerable people such as the elderly and children face an emergency, there is a need to respond quickly and appropriately, but there are still few systems that can achieve this, and existing technologies are insufficient. In particular, there is a lack of means to detect serious abnormalities such as falls and kidnappings early and to take prompt and appropriate action.
[0393] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0394] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information and gyro sensor information to the server, means for analyzing the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, means for activating a security buzzer when an abnormality is detected, means for sending the voice information and gyro sensor information acquired by the terminal to the server via Wi-Fi or a mobile data network, and means for collecting and storing various sensor information at time intervals. This makes it possible to quickly collect and analyze voice information and gyro sensor information and take appropriate action when elderly people, children, or others face an emergency.
[0395] "Means for acquiring audio information" refers to the means by which a terminal uses a built-in microphone or other audio recording device to collect surrounding audio as digital data.
[0396] "Means for acquiring gyro sensor information" refers to the means by which a device uses its built-in gyro sensor to collect data on the user's movements and tilt in real time.
[0397] "Means for transmitting acquired voice information and gyro sensor information to a server" refers to means for transmitting voice information and gyro sensor information collected by the terminal to a server via a communication means such as Wi-Fi or a mobile data network.
[0398] "Means for the server to analyze audio information and gyro sensor information to detect anomalies" refers to means for the server to use machine learning models or other analytical technologies to analyze received audio information and gyro sensor information and detect anomalies.
[0399] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to the means for informing the user and their relatives of an abnormality by sending an email or app notification when the server detects an abnormality.
[0400] The "means for sending a notification to the police when a serious abnormality is detected" is a means for the server to quickly notify the police when it detects a serious abnormality.
[0401] The "means for activating the security buzzer when an abnormality is detected" is a means for the server to send an instruction to the terminal to activate the security buzzer of the terminal when an abnormality is detected.
[0402] "Means for transmitting voice information and gyro sensor information acquired by a terminal to a server via Wi-Fi or a mobile data network" refers to means for transmitting voice information and gyro sensor information collected by a terminal to a server using wireless communication technology.
[0403] "Means for collecting and storing various sensor information at regular intervals" refers to means by which the terminal automatically collects and stores audio information and gyro sensor information at regular intervals.
[0404] This invention is a system for responding quickly and appropriately to emergencies faced by vulnerable people such as the elderly and children. This system acquires voice and gyro sensor information, transmits it to a server for analysis, and activates an alert or a security buzzer when an abnormality is detected.
[0405] Data Acquisition
[0406] The device (smartphone) acquires audio information using a built-in microphone. This is done using a dedicated application on the smartphone. The audio information is collected at regular intervals (for example, every 5 seconds) and saved as data. The device also uses a built-in gyro sensor to monitor the user's movements and tilt in real time. The gyro sensor information is also saved at regular intervals. In this way, the device can acquire audio data and gyro sensor data simultaneously.
[0407] Sending data to the server
[0408] The voice and gyro sensor information collected by the device is sent to a server via Wi-Fi or mobile data network in real time or at regular intervals. The server receives the sent data and stores it in a database. During this process, communication between the device and the server is encrypted to ensure security.
[0409] Anomaly detection
[0410] The server analyzes the received voice information and gyro sensor information. This analysis uses generative AI models and machine learning models. Specifically, natural language processing technology is used to analyze the voice information to detect specific keywords (e.g., "help" or "danger"). Additionally, analysis of the gyro sensor information detects sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration). If an abnormality is detected as a result of the analysis, the system proceeds to the next stage.
[0411] Sending alerts
[0412] If the server detects an abnormality, it evaluates the urgency of that abnormality. If a minor abnormality is detected (for example, if only audio containing specific keywords is detected), the server sends an alert to the user and their relatives. This alert is sent via email or a dedicated application. If a serious abnormality is detected (for example, if audio information and abnormal behavior are detected simultaneously), the server also notifies the police. This allows for a swift emergency response.
[0413] Activating the security alarm
[0414] If the server detects a serious abnormality, it will send a command to the device to activate the security alarm. Activating the security alarm will alert people around the user to the abnormality and make it clear that it is an emergency. This function will enable people around the device to take early action.
[0415] Specific examples
[0416] Falls in the elderly
[0417] If an elderly person suddenly falls, the device's gyro sensor detects the sudden movement and sends that information to the server. The server analyzes the received gyro sensor information and detects the abnormality as a fall. If the server determines that the abnormality is serious, it notifies the family and also reports the incident to the police. The server then sends an instruction to activate the device's security alarm, which then emits a high-volume alarm to alert those around it to the abnormality.
[0418] Child abduction
[0419] In the event of a child abduction, the device's gyro sensor will detect any sudden movements or changes in behavior, and the child's cries for help will be picked up as audio information. This information is sent to a server, which analyzes it to detect any abnormalities. The server will determine this to be a very serious anomaly and immediately notify the police. At the same time, a notification will be sent to the parents, and a security alarm will be activated. This is expected to allow for early response.
[0420] Prompt Sentence Examples
[0421] "Please explain a system that can acquire voice and gyro sensor information and detect abnormalities when an elderly person suddenly collapses."
[0422] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0423] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0424] Step 1: Get the data
[0425] The device uses a built-in microphone to acquire audio information. This is done using a dedicated application, and audio is collected at regular intervals (e.g., every 5 seconds). The input is ambient audio, and digitized audio data is obtained as output. The device also uses a built-in gyro sensor to monitor the user's movement and tilt. This gyro sensor information is also collected at regular intervals, and acceleration data for the x-, y-, and z-axes is obtained as output.
[0426] Step 2: Send data to the server
[0427] The device sends the collected voice information and gyro sensor information to the server via Wi-Fi or mobile data network. The input is the voice data and gyro sensor data collected by the device, and this data is sent to the server as packets at regular time intervals (e.g., 5 seconds). The output is the data packets sent to the server.
[0428] Step 3: Detect anomalies
[0429] The server analyzes the received voice information and gyro sensor information. Generative AI models and machine learning models are used for this analysis. Based on the voice data input, natural language processing techniques are used to detect specific keywords (e.g., "help," "danger"), and based on the gyro sensor data input, sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration) are detected. The analysis results are obtained as output, and a flag is raised if an abnormality is detected.
[0430] Step 4: Sending an alert
[0431] If an anomaly is detected, the server evaluates the anomaly and sends an alert to the user and their relatives. The input is the analysis result, and if the anomaly flag is true, an action is triggered. If the anomaly is minor, the output is an email or app notification sent to the user and their relatives. If a serious anomaly is detected, a notification is also sent to the police.
[0432] Step 5: Activate the security alarm
[0433] If a serious abnormality is detected, the server sends an instruction to the terminal to activate the security buzzer. The input is the abnormality evaluation result, and if the serious abnormality flag is true, an instruction to activate the security buzzer is sent as output to the terminal. Upon receiving this instruction, the terminal activates the security buzzer and emits a high-volume alarm to alert those in the vicinity of the abnormality.
[0434] (Application example 1)
[0435] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0436] In recent years, ensuring the safety of vulnerable people, including the elderly and children, has become an important issue. However, current security systems often lack the ability to recognize abnormalities in real time or respond quickly. Furthermore, few systems that detect abnormalities using both audio and motion information efficiently integrate these data, resulting in problems such as false detection and delayed response. The present invention aims to solve these problems by providing a system that can quickly and accurately detect and respond to abnormalities, especially when the elderly or children are in danger.
[0437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0438] In this invention, the server includes means for analyzing voice information and gyro sensor information to detect abnormalities, means for analyzing specific keywords from the voice information and detecting sudden changes in movement from the gyro sensor information, and means for sending an alert and activating a security buzzer if necessary if an abnormality is detected based on the analysis results. This makes it possible to comprehensively monitor the user's voice and movement, respond quickly in the event of an abnormality, and notify the user and their relatives or those around them of an emergency by activating the security buzzer.
[0439] "Audio information" is data that describes the user's environmental sounds and speech content.
[0440] A "gyro sensor" is a device used to detect the user's movement and tilt.
[0441] The "server" is a computer system that analyzes collected audio information and gyro sensor information, and detects and notifies users of abnormalities.
[0442] An "alert" is a notification sent to the user, their relatives, and, if necessary, the police or other relevant parties when an abnormality is detected.
[0443] A "security buzzer" is a device that emits a sound when an abnormality is detected to alert those around to danger.
[0444] "Analysis" is a process for detecting the presence or absence of abnormalities based on collected audio information and gyro sensor information.
[0445] The "specific keywords" are words and phrases that mainly indicate an emergency in the audio information that indicates an abnormality.
[0446] "Sudden changes in movement" refers to sudden body movements that deviate from normal movement.
[0447] The present invention provides a system that uses voice information and gyro sensor information to monitor the safety of a user and responds quickly when an abnormality is detected. Specific embodiments for carrying out the present invention will be described below.
[0448] System Configuration
[0449] Hardware:
[0450] Smartphone: Equipped with a built-in microphone and gyro sensor, it acquires voice and gyro sensor information.
[0451] Server: Analyzes audio and gyro sensor information to detect abnormalities, send notifications, activate the security alarm, etc.
[0452] software:
[0453] Dedicated application: Equipped with functions to collect voice information and gyro sensor information and send it to a server.
[0454] Cloud server: Uses AWS, GCP, Azure, etc. to analyze and manage data.
[0455] Machine learning model: Using Python or similar software, an algorithm is implemented to analyze audio information and gyro sensor information and detect anomalies.
[0456] Speech recognition technology: Using technologies such as IBM Watson and Google Speech-to-Text, voice data is converted into text and specific keywords are detected.
[0457] Program Processing Overview
[0458] Acquiring audio information
[0459] The user's smartphone uses a built-in microphone to capture voice information, which is then sent to a server at regular intervals via a dedicated application.
[0460] Get gyro sensor information
[0461] The smartphone's built-in gyro sensor is used to monitor the user's movements and tilt in real time, and the application periodically sends the data to a server.
[0462] Data analysis on the server
[0463] The server analyzes the received voice and gyro sensor information. The voice information is converted into text using speech recognition technology to detect whether it contains specific keywords (such as "help" or screams). The gyro sensor information is analyzed using a machine learning model to detect sudden changes in movement (e.g., falling).
[0464] Sending alerts
[0465] If an abnormality is detected, the server evaluates its severity, and if it is a minor abnormality, it sends an app notification to the user and their relatives. If it is a serious abnormality, it also notifies the police.
[0466] Activating the security alarm
[0467] If a serious abnormality is detected, a security alarm will be activated on the smartphone, alerting people nearby.
[0468] Specific examples
[0469] Falls in the elderly
[0470] If an elderly person suddenly falls, the smartphone's gyro sensor will detect the sudden movement and send the information to the server. The server will then detect the fall, notify the family, and notify the police. It will also activate a security alarm to alert people nearby of the emergency.
[0471] Child abduction prevention
[0472] In the case of a child being kidnapped, the smartphone's gyro sensor will detect any sudden changes in movement and capture any voices crying for help. This information will be sent to a server, which will then analyze it and determine that the child has been kidnapped. The server will then immediately notify the police, and a notification will be sent to the parents, who will then activate the security alarm. This is expected to allow for early response.
[0473] Prompt Sentence Examples
[0474] I'd like to see some code examples for analyzing data from an app to detect anomalies in user voice and behavior. The voice data needs to detect specific keywords, such as "help" or "scream." Also, the gyro sensor information needs to detect sudden changes in behavior. I'll be using Python.
[0475] As described above, the system of the present invention can quickly and effectively protect the safety of elderly people and children by using voice information and gyro sensor information.
[0476] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0477] Step 1:
[0478] Acquiring audio information
[0479] The user acquires audio information using a smartphone. The smartphone's built-in microphone collects the user's ambient sounds and speech, and saves the audio data at regular intervals through a dedicated application. The input is the ambient sounds and the user's speech, and the output is the saved audio data.
[0480] Step 2:
[0481] Get gyro sensor information
[0482] The device uses a built-in gyro sensor to monitor the user's movements in real time. The gyro sensor captures acceleration information along the x-, y-, and z-axes, and periodically saves this as data through a dedicated application. The input is the user's movements, and the output is the captured gyro sensor information.
[0483] Step 3:
[0484] Voice information and gyro sensor information sent to the server
[0485] The smartphone transmits the stored voice and gyro sensor information to the server at regular intervals using a secure communication protocol (such as HTTPS). The input is the voice data and gyro sensor information, and the output is the data transmitted to the server.
[0486] Step 4:
[0487] Saving data to the server
[0488] The server stores the received audio information and gyro sensor information in a database. The input is the transmitted data, and the output is the data stored in the database.
[0489] Step 5:
[0490] Analysis of audio information
[0491] The server uses speech recognition technology to convert the stored voice information into text, and analyzes this text data to see if it contains specific keywords (such as "help" or a cry). The input is the stored voice information, and the output is the text data and the analysis results.
[0492] Step 6:
[0493] Analysis of gyro sensor information
[0494] The server analyzes the gyro sensor information using a machine learning model, detects sudden changes in movement (e.g., falling or sudden acceleration), and evaluates the user's situation. The input is the stored gyro sensor information, and the output is the analysis result.
[0495] Step 7:
[0496] Anomaly detection
[0497] The server detects abnormalities based on the results of analyzing voice information and gyro sensor information. It evaluates the degree of abnormality based on the detection results and decides how to respond. The input is the results of voice analysis and gyro analysis, and the output is the presence or absence of an abnormality and its degree.
[0498] Step 8:
[0499] Sending alerts
[0500] The server sends an alert to the user and their relatives depending on the detected anomaly. If the anomaly is minor, it sends an app notification, and if it is severe, it also notifies the police. The input is the degree of anomaly, and the output is the alert sent.
[0501] Step 9:
[0502] Activating the security alarm
[0503] If the server determines that a serious abnormality exists, it sends a command to the smartphone to activate the security buzzer. This notifies people around the smartphone of the emergency. The input is the degree of abnormality, and the output is the activation of the security buzzer.
[0504] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0505] The present invention is a system that includes a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected, and further includes an emotion engine that recognizes the user's emotions.
[0506] Overview of program processing
[0507] 1. Data Acquisition
[0508] The device (smartphone) acquires voice information using a built-in microphone. The voice information is collected at regular intervals and temporarily stored within the device. Motion data is also acquired using a gyro sensor and is also stored within the device.
[0509] 2. Emotion recognition
[0510] The device transmits the acquired voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with other sensor data.
[0511] 3. Sending data to the server
[0512] The device sends collected voice information, gyro sensor information, and emotion information to a server, which sends this data in real time or at specified intervals.
[0513] 4. Anomaly Detection
[0514] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models to detect whether specific keywords (e.g., "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data. It also evaluates whether the user is under high stress based on the emotion information provided by the emotion engine.
[0515] 5. Sending alerts
[0516] If an anomaly is detected, the server will trigger an appropriate alert depending on the severity: if it is a minor anomaly, an alert will be sent to the user and their relatives; if it is a serious anomaly, a notification will also be sent to the police.
[0517] 6. Activating the security alarm
[0518] If a serious abnormality is detected, the server sends a command to the device to activate the burglar alarm, which then automatically sounds the alarm to alert people in the vicinity.
[0519] Specific examples
[0520] For example, let's simulate an elderly person suddenly collapsing at home. The device's gyro sensor detects the sudden movement and sends gyro data to the server as a falling motion. At the same time, a voice calling for help is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server, and the analysis results indicate a serious abnormality. The server notifies the family and also alerts the police. It also sends instructions to the device to activate the security alarm. This allows people in the vicinity to be quickly notified of the abnormality.
[0521] The system also simulates a situation where a child is about to be kidnapped in a park. In this case, the device detects sudden movements and the voice calling for help, while the emotion engine detects strong fear. This information is sent to the server, which determines it as a serious abnormality. The server then immediately notifies the police and parents and instructs the device to activate the security alarm, enabling early action to be taken and helping to ensure the child's safety.
[0522] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children by integrating voice information, gyro sensor information, and emotional information.
[0523] The processing flow will be explained below.
[0524] Step 1:
[0525] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0526] Step 2:
[0527] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0528] Step 3:
[0529] The device sends the collected voice information to the emotion engine, which analyzes the voice data in real time and recognizes the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is temporarily stored in the device along with other sensor data.
[0530] Step 4:
[0531] The device transmits voice information, gyro sensor information, and emotion information to the server. This data is sent to the server in real time or at specified intervals. The API used for transmission is encrypted as a security measure.
[0532] Step 5:
[0533] The server stores the received voice, gyro sensor, and emotion information in a database, which can then be compared with past data for analysis to detect anomalies.
[0534] Step 6:
[0535] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[0536] Step 7:
[0537] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[0538] Step 8:
[0539] The server analyzes the emotional information and evaluates whether the user is under high stress or fear based on the data provided by the emotion engine.
[0540] Step 9:
[0541] The server uses the results of voice, gyro sensor, and emotion analysis to determine whether an anomaly has been detected, and if so, evaluates whether the anomaly is minor or major.
[0542] Step 10:
[0543] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[0544] Step 11:
[0545] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[0546] Step 12:
[0547] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[0548] Step 13:
[0549] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[0550] Example 2
[0551] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0552] Ensuring the safety of elderly people and children in modern society is a critical issue, requiring rapid and accurate responses, especially in emergencies. Conventional anomaly detection systems respond based solely on voice and motion information, and do not take into account changes in emotions. Such systems have difficulty determining the level of urgency, and there is a risk of delays in issuing appropriate alerts or notifications. Therefore, in order to more reliably protect the personal safety of elderly people and children, there is a need to develop a system that can also comprehensively analyze emotional information.
[0553] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0554] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information to an emotion engine to analyze the user's emotion, means for transmitting the analyzed emotion information to the server, means for transmitting the voice information and gyro sensor information to the server, means for the server to analyze the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, and means for activating a security buzzer when an abnormality is detected. This makes it possible to comprehensively analyze the user's emotion information, appropriately determine which abnormalities are of high urgency, and respond quickly.
[0555] The "means for acquiring voice information" is a means for collecting voice data using a built-in microphone or other voice input device.
[0556] "Means for acquiring gyro sensor information" refers to a means for using a gyro sensor to measure fluctuations in the user's movements and posture and collect that data.
[0557] The "means for transmitting acquired voice information and gyro sensor information to a server" refers to a means having the function of transmitting collected voice data and gyro data to a server via wireless communication or cable connection.
[0558] "Means for the server to analyze audio information and gyro sensor information to detect abnormalities" refers to means for analyzing audio data and gyro data within the server and executing algorithms or programs that detect abnormalities based on the results.
[0559] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to a means that has the function of sending a notification to the user and their pre-designated relatives when an abnormality is detected.
[0560] "Means for sending a notification to the police when a serious abnormality is detected" refers to means that has the function of sending an emergency notification to public institutions such as the police when a highly urgent abnormality is detected.
[0561] "Means for activating the security alarm when an abnormality is detected" refers to means having the function of sounding the security alarm of the terminal to notify those around when the system detects an abnormality.
[0562] An "emotion engine that recognizes user emotions" is an algorithm or program that analyzes acquired voice data to identify and identify the user's emotional state.
[0563] The "means for transmitting acquired voice information to an emotion engine and analyzing the user's emotions" refers to a means having a function for transmitting collected voice data to an emotion engine and analyzing the emotional state.
[0564] The "means for transmitting analyzed emotion information to a server" is a means having a function for transmitting emotion data analyzed by an emotion engine to a server.
[0565] The system of the present invention collects user voice and gyro sensor information, sends it to a server for analysis, detects abnormalities, and issues appropriate alerts and notifications. It also has the ability to collect and analyze user emotional information to more accurately determine the urgency of the abnormality.
[0566] Hardware and software used
[0567] This system uses the following hardware and software:
[0568] Hardware:
[0569] 1. Device (smartphone): Collects voice information using the built-in microphone and acquires movement data using the built-in gyro sensor.
[0570] 2. Server: Analyzes data, detects anomalies, and manages alerts and notifications.
[0571] software:
[0572] 1. Emotion engine: Software for emotion recognition (e.g., IBM Watson or Google Cloud Speech-to-Text).
[0573] 2. Machine learning model: Anomaly detection model trained using TensorFlow and PyTorch.
[0574] 3. Data transmission and management application: An application that manages data transmission from the terminal to the server and instructions from the server.
[0575] Data Acquisition and Processing
[0576] The device uses a built-in microphone to collect audio information at regular intervals (e.g., every 5 seconds). The collected audio data is temporarily stored in the device. Similarly, the built-in gyro sensor is used to obtain motion data, which is also stored in the device.
[0577] The device then sends the collected voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with the collected gyro data.
[0578] The server uses voice analysis algorithms and machine learning models to analyze the received voice information, gyro data, and emotional information. If an abnormality is detected, the server triggers an appropriate alert depending on the urgency. If the abnormality is minor, the server notifies the user and their relatives, and if it is serious, it notifies the police. If a serious abnormality is detected, the server sends an instruction to the device to sound the security alarm.
[0579] Specific examples
[0580] Example 1: An elderly person suddenly collapses at home
[0581] The device's gyro sensor detects violent movements and sends gyro data to the server as a falling motion. At the same time, voice information calling for "help" is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server in real time. The server analyzes this information using a voice analysis algorithm and machine learning model and determines that it is a serious abnormality. The server then sends a notification to family members and alerts the police. It can also send instructions to the device to activate a security alarm, quickly alerting people in the vicinity of an abnormality.
[0582] Example 2: A child is nearly kidnapped in a park
[0583] The device detects sudden movements and cries for help, while the emotion engine detects strong fear. This information is sent to the server in real time. The server immediately detects any abnormalities and notifies the police and parents as serious problems. The server also sends instructions to the device to sound a security alarm, alerting people in the vicinity. This encourages rapid response and helps ensure the safety of children.
[0584] Prompt Sentence Examples
[0585] Example prompt: "Write an outline of an emergency alert system for people with disabilities. The system has voice recognition, emotion detection, and anomaly detection capabilities, and the ability to notify users of detected anomalies."
[0586] Sample prompt: "Describe an anomaly detection system for when an elderly person suddenly collapses. The system uses voice, gyro sensor, and emotion data to detect the anomaly."
[0587] In this way, the system of the present invention makes it possible to quickly and effectively protect the safety of elderly people and children by integrating voice information, gyro sensor information, and emotion information.
[0588] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0589] Step 1:
[0590] Data Acquisition
[0591] The device (smartphone) acquires audio information using a built-in microphone. The microphone continuously captures audio and collects data at regular intervals (for example, every 5 seconds). The input is the user's voice, and the output is audio data stored in the device. Similarly, the device's built-in gyro sensor captures the user's movement data in real time. Gyro data is also collected at regular intervals and stored in the device. The input is the user's movement, and the output is gyro sensor data.
[0592] Step 2:
[0593] emotion recognition
[0594] The device sends the acquired voice information in real time to an emotion engine (for example, IBM Watson or Google Cloud Speech-to-Text) to recognize the user's emotions. The input is the voice information stored on the device, and the output is the emotional state (anger, anxiety, sadness, etc.) obtained from the emotion engine. For example, if the user is angry, the emotion engine will recognize "anger" based on the tone and content of the voice. This emotion information is temporarily stored on the device and sent to the server along with gyro data.
[0595] Step 3:
[0596] Sending data to the server
[0597] The device sends voice information, gyro sensor information, and emotion information to the server. The input is the voice data, gyro data, and emotion information stored on the device, and the output is each data sent to the server. This data transmission is done in real time, but depending on the communication environment, it may also be sent at specified time intervals (for example, every minute).
[0598] Step 4:
[0599] Anomaly detection
[0600] The server analyzes the received voice information, gyro sensor information, and emotional information. The input is the voice data, gyro data, and emotional information sent to the server, and the output is the anomaly detection results. The server is equipped with a machine learning model (for example, a model trained using TensorFlow or PyTorch) and uses this to detect anomalies. Specifically, the server uses a voice analysis algorithm to detect whether the voice data contains specific keywords (for example, "help" or screams). It also analyzes patterns in the gyro data to determine whether the user has suddenly fallen. Furthermore, it evaluates whether the user is experiencing strong stress or fear based on the emotional information provided by the emotion engine.
[0601] Step 5:
[0602] Sending alerts
[0603] If an anomaly is detected, the server triggers an alert depending on its urgency. The input is the anomaly detection result, and the output is the alert to be sent. If the anomaly is minor (for example, a small noise or a slight fall), the server will send an SMS or app notification to the user and their relatives. If the anomaly is serious (for example, if a cry for help or strong emotional expression is detected), the server will immediately notify the police. The notification also includes location information.
[0604] Step 6:
[0605] Activating the security alarm
[0606] When a serious abnormality is detected, the server sends a command to the device to activate the security buzzer. The input is the serious abnormality detection result, and the output is the sound of the security buzzer. The device receives this command and sounds the security buzzer using its built-in speaker. This alerts people in the vicinity to the abnormality and encourages them to take early action.
[0607] (Application example 2)
[0608] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0609] Conventional security systems rely solely on audio and motion information to detect abnormalities, resulting in low accuracy in detecting abnormalities and often making it difficult to determine whether an abnormality has occurred because they do not take into account the emotional state of the user. Furthermore, even if an abnormality is detected, prompt action may not be possible, posing a challenge to ensuring user safety.
[0610] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0611] The smart glasses include a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, a means for activating a security buzzer when an abnormality is detected, a means including an emotion engine that recognizes the user's emotions, a means for transmitting emotion information recognized by the emotion engine to a server, and a means for supporting the safety of users using the smart glasses. This improves the accuracy of abnormality detection and enables prompt and appropriate responses.
[0612] A "means for acquiring voice information" is a device or mechanism that detects a user's voice and records and collects it as digital data.
[0613] "Means for acquiring gyro sensor information" refers to devices or mechanisms that measure changes in a user's movements and posture, and record and collect them as digital data.
[0614] The "means for transmitting to a server" refers to a device or method for transmitting the acquired voice information and gyro sensor information to a server via a communication means such as wireless communication or the Internet.
[0615] "Means for analyzing and detecting anomalies" refers to a processing system that analyzes the data received by the server using algorithms, machine learning models, etc., to identify anomalies.
[0616] An "alert sending means" is a device or method for sending a warning or caution notification to a designated recipient when an abnormality is detected.
[0617] "Means for sending notifications" refers to devices or methods for promptly contacting the user's relatives, police, or other relevant authorities in the event of a serious abnormality.
[0618] "Means for activating a security alarm" refers to the mechanism or control method for activating a physical audio alarm device (security alarm) when an abnormality is detected.
[0619] An "emotion engine" is an algorithm or software that analyzes acquired voice information and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.).
[0620] "Smart glasses" are high-performance glasses that can acquire audio information and gyro sensor information, and can also recognize the user's emotions using an emotion engine, displaying visual information and communicating.
[0621] In this invention, we will build a system in which smart glasses and other applicable devices acquire voice information and gyro sensor information and send it to a server. This allows for quick and appropriate response when an abnormality occurs. Specific implementation methods are shown below.
[0622] First, we use smart glasses as the hardware. The smart glasses have built-in microphones and gyro sensors, which enable them to collect voice and motion information. They also have wireless communication capabilities such as Wi-Fi and Bluetooth, which allow them to send collected data to a server. Second, we use Google Cloud's Speech-to-Text API and Emotion API as emotion engines for emotion recognition.
[0623] The data received by the server is:
[0624] 1. Audio information: The user's voice is picked up by a microphone and recorded as digital data.
[0625] 2. Gyro sensor information: Detects user movements and collects their movement data.
[0626] 3. Emotion information: Emotion data analyzed by the emotion engine.
[0627] The server analyzes the received data and detects anomalies using machine learning models (for example, TensorFlow or PyTorch). Based on the analysis results, it determines the urgency of the anomaly and sends an appropriate alert. Specifically, it notifies the user and their relatives using Twilio's SMS API or Push Notification API.
[0628] If the abnormality is determined to be serious, the server sends a notification to the police and instructs the device to activate a security alarm, which activates a physical audio alarm and alerts people in the vicinity to the abnormality.
[0629] As a concrete example, consider the case where an elderly person suddenly falls while out. In this case, the gyro sensor in the smart glasses detects the fall, and the microphone picks up the cry of "Ouch!" The emotion engine distinguishes between strong pain and fear, and sends this data to the server. The server determines this to be a serious abnormality, sends an emergency call to relatives, and activates the security alarm.
[0630] Prompt Sentence Examples
[0631] "Please build a system that collects voice and gyro sensor information in real time to detect when the user falls or has a sudden change in behavior, and sends the data along with the user's emotional state to a server for analysis and alerts when an abnormality occurs."
[0632] In this way, the invention provides a system that uses smart glasses to collect voice information and gyro sensor information, and also integrates and analyzes emotional information, thereby improving the accuracy of anomaly detection and enabling prompt and appropriate responses.
[0633] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0634] Step 1: Acquire audio and gyro sensor information
[0635] The device (smart glasses) collects the user's voice information using a built-in microphone. At the same time, it acquires the user's movement data using a gyro sensor. The acquired voice information and gyro sensor information are temporarily stored in the device. This input data is the raw data for monitoring the user's voice and movement in real time.
[0636] Step 2: Emotion Recognition
[0637] The device sends the collected voice information to an emotion engine in real time to analyze the user's emotions. The emotion engine (for example, Google Cloud's Speech-to-Text API or Emotion API) analyzes the voice data and identifies the emotional state (anger, anxiety, sadness, etc.). It receives voice information as input data and outputs emotional information. This emotional information is also temporarily stored on the device.
[0638] Step 3: Send data
[0639] The device transmits the acquired voice information, gyro sensor information, and emotion information to the server via Wi-Fi or Bluetooth. This data is sent to the server as input data for analysis.
[0640] Step 4: Anomaly detection
[0641] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models (TensorFlow and PyTorch) to detect whether specific keywords (such as "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the data patterns from the gyro sensor. It also evaluates the user's level of stress based on the emotion information provided by the emotion engine. The analysis results indicate whether there is an abnormality and its urgency.
[0642] Step 5: Sending an alert
[0643] If an abnormality is detected, an alert is sent according to the urgency. If the abnormality is minor, the server sends an alert to the user and their relatives. SMS API and Push Notification API are used for sending the alert. If a serious abnormality occurs, the server also sends a notification to the police. The input data is the analysis result, and the output data is the alert notification.
[0644] Step 6: Activate the security alarm
[0645] If a serious abnormality is detected, the server sends an instruction to the device to activate the security alarm. This activates the security alarm in the smart glasses and alerts people nearby. The input data is the instruction from the server, and the output data is the activation of the security alarm.
[0646] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0647] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0648] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0649] [Third embodiment]
[0650] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0651] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0652] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0653] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0654] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0655] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0656] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0657] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0658] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0659] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0660] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0661] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0662] The present invention is a system that includes a means for acquiring audio information, a means for acquiring gyro sensor information, a means for transmitting the acquired audio information and gyro sensor information to a server, a means for the server to analyze the audio information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected.
[0663] Overview of program processing
[0664] 1. Data Acquisition
[0665] The device (smartphone) acquires voice information using a built-in microphone. This is done using a dedicated application on the device. The voice information is collected at regular intervals and saved as data. The device also acquires information from its built-in gyro sensor in real time to monitor the user's movements and tilt. This gyro sensor information is also saved at regular intervals.
[0666] 2. Sending data to the server
[0667] The voice and gyro sensor information collected by the device is sent to a server in real time or at regular intervals, and the server receives the data and stores it in a database.
[0668] 3. Anomaly detection
[0669] The server analyzes audio and gyro sensor information to detect abnormalities. This analysis is performed using a machine learning model. Audio analysis detects specific keywords (such as screams or "help me"). Gyro sensor analysis also detects sudden movements and unnatural fluctuations (such as falling or sudden changes in acceleration).
[0670] 4. Sending alerts
[0671] If an abnormality is detected, the server evaluates it, and if it is a minor abnormality, it sends an alert to the user and their relatives. This alert is sent via email or app notification. If a serious abnormality is detected, the server also notifies the police, allowing for a prompt emergency response.
[0672] 5. Activating the security alarm
[0673] If an abnormality is detected and is deemed to be particularly serious, the server sends a command to the device to activate the security alarm, thereby notifying people around the user of the abnormality and alerting them to the emergency.
[0674] Specific examples
[0675] Below is a concrete example of how the system of the present invention actually works.
[0676] Consider the case where an elderly person suddenly collapses. At this time, the gyro sensor on the device detects the sudden movement and sends the gyro sensor information to the server as a falling motion. The server analyzes the received gyro information and detects an abnormality such as a fall. If the server determines that this abnormality is serious, it notifies the family and also reports the incident to the police. The server also sends instructions to the device to activate the security alarm. This alerts people in the vicinity to the abnormality, enabling a swift response.
[0677] The system also simulates a child being abducted. In this case, the gyro sensor detects any sudden movements or changes in behavior, and the child's cries for help are captured as audio information. This information is sent to the server, which analyzes it to detect any abnormalities. The server determines this to be a very serious anomaly and immediately notifies the police. At the same time, a notification is sent to the parents, and a security alarm is activated. This is expected to allow for early response.
[0678] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0679] The processing flow will be explained below.
[0680] Step 1:
[0681] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0682] Step 2:
[0683] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0684] Step 3:
[0685] The device sends the collected voice and gyro sensor information to a server. In particular, this data is sent to the server in real time or at specified intervals via internet communication. The API used for transmission is encrypted as a security measure.
[0686] Step 4:
[0687] The server stores the received audio and gyro sensor information in a database, which allows analysis to detect anomalies by comparing it with past data.
[0688] Step 5:
[0689] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[0690] Step 6:
[0691] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[0692] Step 7:
[0693] The server determines whether an abnormality is detected based on the results of voice analysis and gyro sensor analysis, and if so, evaluates whether the abnormality is minor or serious.
[0694] Step 8:
[0695] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[0696] Step 9:
[0697] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[0698] Step 10:
[0699] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[0700] Step 11:
[0701] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[0702] Example 1
[0703] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0704] In modern society, when vulnerable people such as the elderly and children face an emergency, there is a need to respond quickly and appropriately, but there are still few systems that can achieve this, and existing technologies are insufficient. In particular, there is a lack of means to detect serious abnormalities such as falls and kidnappings early and to take prompt and appropriate action.
[0705] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0706] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information and gyro sensor information to the server, means for analyzing the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, means for activating a security buzzer when an abnormality is detected, means for sending the voice information and gyro sensor information acquired by the terminal to the server via Wi-Fi or a mobile data network, and means for collecting and storing various sensor information at time intervals. This makes it possible to quickly collect and analyze voice information and gyro sensor information and take appropriate action when elderly people, children, or others face an emergency.
[0707] "Means for acquiring audio information" refers to the means by which a terminal uses a built-in microphone or other audio recording device to collect surrounding audio as digital data.
[0708] "Means for acquiring gyro sensor information" refers to the means by which a device uses its built-in gyro sensor to collect data on the user's movements and tilt in real time.
[0709] "Means for transmitting acquired voice information and gyro sensor information to a server" refers to means for transmitting voice information and gyro sensor information collected by the terminal to a server via a communication means such as Wi-Fi or a mobile data network.
[0710] "Means for the server to analyze audio information and gyro sensor information to detect anomalies" refers to means for the server to use machine learning models or other analytical technologies to analyze received audio information and gyro sensor information and detect anomalies.
[0711] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to the means for informing the user and their relatives of an abnormality by sending an email or app notification when the server detects an abnormality.
[0712] The "means for sending a notification to the police when a serious abnormality is detected" is a means for the server to quickly notify the police when it detects a serious abnormality.
[0713] The "means for activating the security buzzer when an abnormality is detected" is a means for the server to send an instruction to the terminal to activate the security buzzer of the terminal when an abnormality is detected.
[0714] "Means for transmitting voice information and gyro sensor information acquired by a terminal to a server via Wi-Fi or a mobile data network" refers to means for transmitting voice information and gyro sensor information collected by a terminal to a server using wireless communication technology.
[0715] "Means for collecting and storing various sensor information at regular intervals" refers to means by which the terminal automatically collects and stores audio information and gyro sensor information at regular intervals.
[0716] This invention is a system for responding quickly and appropriately to emergencies faced by vulnerable people such as the elderly and children. This system acquires voice and gyro sensor information, transmits it to a server for analysis, and activates an alert or a security buzzer when an abnormality is detected.
[0717] Data Acquisition
[0718] The device (smartphone) acquires audio information using a built-in microphone. This is done using a dedicated application on the smartphone. The audio information is collected at regular intervals (for example, every 5 seconds) and saved as data. The device also uses a built-in gyro sensor to monitor the user's movements and tilt in real time. The gyro sensor information is also saved at regular intervals. In this way, the device can acquire audio data and gyro sensor data simultaneously.
[0719] Sending data to the server
[0720] The voice and gyro sensor information collected by the device is sent to a server via Wi-Fi or mobile data network in real time or at regular intervals. The server receives the sent data and stores it in a database. During this process, communication between the device and the server is encrypted to ensure security.
[0721] Anomaly detection
[0722] The server analyzes the received voice information and gyro sensor information. This analysis uses generative AI models and machine learning models. Specifically, natural language processing technology is used to analyze the voice information to detect specific keywords (e.g., "help" or "danger"). Additionally, analysis of the gyro sensor information detects sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration). If an abnormality is detected as a result of the analysis, the system proceeds to the next stage.
[0723] Sending alerts
[0724] If the server detects an abnormality, it evaluates the urgency of that abnormality. If a minor abnormality is detected (for example, if only audio containing specific keywords is detected), the server sends an alert to the user and their relatives. This alert is sent via email or a dedicated application. If a serious abnormality is detected (for example, if audio information and abnormal behavior are detected simultaneously), the server also notifies the police. This allows for a swift emergency response.
[0725] Activating the security alarm
[0726] If the server detects a serious abnormality, it will send a command to the device to activate the security alarm. Activating the security alarm will alert people around the user to the abnormality and make it clear that it is an emergency. This function will enable people around the device to take early action.
[0727] Specific examples
[0728] Falls in the elderly
[0729] If an elderly person suddenly falls, the device's gyro sensor detects the sudden movement and sends that information to the server. The server analyzes the received gyro sensor information and detects the abnormality as a fall. If the server determines that the abnormality is serious, it notifies the family and also reports the incident to the police. The server then sends an instruction to activate the device's security alarm, which then emits a high-volume alarm to alert those around it to the abnormality.
[0730] Child abduction
[0731] In the event of a child abduction, the device's gyro sensor will detect any sudden movements or changes in behavior, and the child's cries for help will be picked up as audio information. This information is sent to a server, which analyzes it to detect any abnormalities. The server will determine this to be a very serious anomaly and immediately notify the police. At the same time, a notification will be sent to the parents, and a security alarm will be activated. This is expected to allow for early response.
[0732] Prompt Sentence Examples
[0733] "Please explain a system that can acquire voice and gyro sensor information and detect abnormalities when an elderly person suddenly collapses."
[0734] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0735] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0736] Step 1: Get the data
[0737] The device uses a built-in microphone to acquire audio information. This is done using a dedicated application, and audio is collected at regular intervals (e.g., every 5 seconds). The input is ambient audio, and digitized audio data is obtained as output. The device also uses a built-in gyro sensor to monitor the user's movement and tilt. This gyro sensor information is also collected at regular intervals, and acceleration data for the x-, y-, and z-axes is obtained as output.
[0738] Step 2: Send data to the server
[0739] The device sends the collected voice information and gyro sensor information to the server via Wi-Fi or mobile data network. The input is the voice data and gyro sensor data collected by the device, and this data is sent to the server as packets at regular time intervals (e.g., 5 seconds). The output is the data packets sent to the server.
[0740] Step 3: Detect anomalies
[0741] The server analyzes the received voice information and gyro sensor information. Generative AI models and machine learning models are used for this analysis. Based on the voice data input, natural language processing techniques are used to detect specific keywords (e.g., "help," "danger"), and based on the gyro sensor data input, sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration) are detected. The analysis results are obtained as output, and a flag is raised if an abnormality is detected.
[0742] Step 4: Sending an alert
[0743] If an anomaly is detected, the server evaluates the anomaly and sends an alert to the user and their relatives. The input is the analysis result, and if the anomaly flag is true, an action is triggered. If the anomaly is minor, the output is an email or app notification sent to the user and their relatives. If a serious anomaly is detected, a notification is also sent to the police.
[0744] Step 5: Activate the security alarm
[0745] If a serious abnormality is detected, the server sends an instruction to the terminal to activate the security buzzer. The input is the abnormality evaluation result, and if the serious abnormality flag is true, an instruction to activate the security buzzer is sent as output to the terminal. Upon receiving this instruction, the terminal activates the security buzzer and emits a high-volume alarm to alert those in the vicinity of the abnormality.
[0746] (Application example 1)
[0747] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0748] In recent years, ensuring the safety of vulnerable people, including the elderly and children, has become an important issue. However, current security systems often lack the ability to recognize abnormalities in real time or respond quickly. Furthermore, few systems that detect abnormalities using both audio and motion information efficiently integrate these data, resulting in problems such as false detection and delayed response. The present invention aims to solve these problems by providing a system that can quickly and accurately detect and respond to abnormalities, especially when the elderly or children are in danger.
[0749] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0750] In this invention, the server includes means for analyzing voice information and gyro sensor information to detect abnormalities, means for analyzing specific keywords from the voice information and detecting sudden changes in movement from the gyro sensor information, and means for sending an alert and activating a security buzzer if necessary if an abnormality is detected based on the analysis results. This makes it possible to comprehensively monitor the user's voice and movement, respond quickly in the event of an abnormality, and notify the user and their relatives or those around them of an emergency by activating the security buzzer.
[0751] "Audio information" is data that describes the user's environmental sounds and speech content.
[0752] A "gyro sensor" is a device used to detect the user's movement and tilt.
[0753] The "server" is a computer system that analyzes collected audio information and gyro sensor information, and detects and notifies users of abnormalities.
[0754] An "alert" is a notification sent to the user, their relatives, and, if necessary, the police or other relevant parties when an abnormality is detected.
[0755] A "security buzzer" is a device that emits a sound when an abnormality is detected to alert those around to danger.
[0756] "Analysis" is a process for detecting the presence or absence of abnormalities based on collected audio information and gyro sensor information.
[0757] The "specific keywords" are words and phrases that mainly indicate an emergency in the audio information that indicates an abnormality.
[0758] "Sudden changes in movement" refers to sudden body movements that deviate from normal movement.
[0759] The present invention provides a system that uses voice information and gyro sensor information to monitor the safety of a user and responds quickly when an abnormality is detected. Specific embodiments for carrying out the present invention will be described below.
[0760] System Configuration
[0761] Hardware:
[0762] Smartphone: Equipped with a built-in microphone and gyro sensor, it acquires voice and gyro sensor information.
[0763] Server: Analyzes audio and gyro sensor information to detect abnormalities, send notifications, activate the security alarm, etc.
[0764] software:
[0765] Dedicated application: Equipped with functions to collect voice information and gyro sensor information and send it to a server.
[0766] Cloud server: Uses AWS, GCP, Azure, etc. to analyze and manage data.
[0767] Machine learning model: Using Python or similar software, an algorithm is implemented to analyze audio information and gyro sensor information and detect anomalies.
[0768] Speech recognition technology: Using technologies such as IBM Watson and Google Speech-to-Text, voice data is converted into text and specific keywords are detected.
[0769] Program Processing Overview
[0770] Acquiring audio information
[0771] The user's smartphone uses a built-in microphone to capture voice information, which is then sent to a server at regular intervals via a dedicated application.
[0772] Get gyro sensor information
[0773] The smartphone's built-in gyro sensor is used to monitor the user's movements and tilt in real time, and the application periodically sends the data to a server.
[0774] Data analysis on the server
[0775] The server analyzes the received voice and gyro sensor information. The voice information is converted into text using speech recognition technology to detect whether it contains specific keywords (such as "help" or screams). The gyro sensor information is analyzed using a machine learning model to detect sudden changes in movement (e.g., falling).
[0776] Sending alerts
[0777] If an abnormality is detected, the server evaluates its severity, and if it is a minor abnormality, it sends an app notification to the user and their relatives. If it is a serious abnormality, it also notifies the police.
[0778] Activating the security alarm
[0779] If a serious abnormality is detected, a security alarm will be activated on the smartphone, alerting people nearby.
[0780] Specific examples
[0781] Falls in the elderly
[0782] If an elderly person suddenly falls, the smartphone's gyro sensor will detect the sudden movement and send the information to the server. The server will then detect the fall, notify the family, and notify the police. It will also activate a security alarm to alert people nearby of the emergency.
[0783] Child abduction prevention
[0784] In the case of a child being kidnapped, the smartphone's gyro sensor will detect any sudden changes in movement and capture any voices crying for help. This information will be sent to a server, which will then analyze it and determine that the child has been kidnapped. The server will then immediately notify the police, and a notification will be sent to the parents, who will then activate the security alarm. This is expected to allow for early response.
[0785] Prompt Sentence Examples
[0786] I'd like to see some code examples for analyzing data from an app to detect anomalies in user voice and behavior. The voice data needs to detect specific keywords, such as "help" or "scream." Also, the gyro sensor information needs to detect sudden changes in behavior. I'll be using Python.
[0787] As described above, the system of the present invention can quickly and effectively protect the safety of elderly people and children by using voice information and gyro sensor information.
[0788] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0789] Step 1:
[0790] Acquiring audio information
[0791] The user acquires audio information using a smartphone. The smartphone's built-in microphone collects the user's ambient sounds and speech, and saves the audio data at regular intervals through a dedicated application. The input is the ambient sounds and the user's speech, and the output is the saved audio data.
[0792] Step 2:
[0793] Get gyro sensor information
[0794] The device uses a built-in gyro sensor to monitor the user's movements in real time. The gyro sensor captures acceleration information along the x-, y-, and z-axes, and periodically saves this as data through a dedicated application. The input is the user's movements, and the output is the captured gyro sensor information.
[0795] Step 3:
[0796] Voice information and gyro sensor information sent to the server
[0797] The smartphone transmits the stored voice and gyro sensor information to the server at regular intervals using a secure communication protocol (such as HTTPS). The input is the voice data and gyro sensor information, and the output is the data transmitted to the server.
[0798] Step 4:
[0799] Saving data to the server
[0800] The server stores the received audio information and gyro sensor information in a database. The input is the transmitted data, and the output is the data stored in the database.
[0801] Step 5:
[0802] Analysis of audio information
[0803] The server uses speech recognition technology to convert the stored voice information into text, and analyzes this text data to see if it contains specific keywords (such as "help" or a cry). The input is the stored voice information, and the output is the text data and the analysis results.
[0804] Step 6:
[0805] Analysis of gyro sensor information
[0806] The server analyzes the gyro sensor information using a machine learning model, detects sudden changes in movement (e.g., falling or sudden acceleration), and evaluates the user's situation. The input is the stored gyro sensor information, and the output is the analysis result.
[0807] Step 7:
[0808] Anomaly detection
[0809] The server detects abnormalities based on the results of analyzing voice information and gyro sensor information. It evaluates the degree of abnormality based on the detection results and decides how to respond. The input is the results of voice analysis and gyro analysis, and the output is the presence or absence of an abnormality and its degree.
[0810] Step 8:
[0811] Sending alerts
[0812] The server sends an alert to the user and their relatives depending on the detected anomaly. If the anomaly is minor, it sends an app notification, and if it is severe, it also notifies the police. The input is the degree of anomaly, and the output is the alert sent.
[0813] Step 9:
[0814] Activating the security alarm
[0815] If the server determines that a serious abnormality exists, it sends a command to the smartphone to activate the security buzzer. This notifies people around the smartphone of the emergency. The input is the degree of abnormality, and the output is the activation of the security buzzer.
[0816] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0817] The present invention is a system that includes a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected, and further includes an emotion engine that recognizes the user's emotions.
[0818] Overview of program processing
[0819] 1. Data Acquisition
[0820] The device (smartphone) acquires voice information using a built-in microphone. The voice information is collected at regular intervals and temporarily stored within the device. Motion data is also acquired using a gyro sensor and is also stored within the device.
[0821] 2. Emotion recognition
[0822] The device transmits the acquired voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with other sensor data.
[0823] 3. Sending data to the server
[0824] The device sends collected voice information, gyro sensor information, and emotion information to a server, which sends this data in real time or at specified intervals.
[0825] 4. Anomaly Detection
[0826] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models to detect whether specific keywords (e.g., "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data. It also evaluates whether the user is under high stress based on the emotion information provided by the emotion engine.
[0827] 5. Sending alerts
[0828] If an anomaly is detected, the server will trigger an appropriate alert depending on the severity: if it is a minor anomaly, an alert will be sent to the user and their relatives; if it is a serious anomaly, a notification will also be sent to the police.
[0829] 6. Activating the security alarm
[0830] If a serious abnormality is detected, the server sends a command to the device to activate the burglar alarm, which then automatically sounds the alarm to alert people in the vicinity.
[0831] Specific examples
[0832] For example, let's simulate an elderly person suddenly collapsing at home. The device's gyro sensor detects the sudden movement and sends gyro data to the server as a falling motion. At the same time, a voice calling for help is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server, and the analysis results indicate a serious abnormality. The server notifies the family and also alerts the police. It also sends instructions to the device to activate the security alarm. This allows people in the vicinity to be quickly notified of the abnormality.
[0833] The system also simulates a situation where a child is about to be kidnapped in a park. In this case, the device detects sudden movements and the voice calling for help, while the emotion engine detects strong fear. This information is sent to the server, which determines it as a serious abnormality. The server then immediately notifies the police and parents and instructs the device to activate the security alarm, enabling early action to be taken and helping to ensure the child's safety.
[0834] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children by integrating voice information, gyro sensor information, and emotional information.
[0835] The processing flow will be explained below.
[0836] Step 1:
[0837] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0838] Step 2:
[0839] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0840] Step 3:
[0841] The device sends the collected voice information to the emotion engine, which analyzes the voice data in real time and recognizes the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is temporarily stored in the device along with other sensor data.
[0842] Step 4:
[0843] The device transmits voice information, gyro sensor information, and emotion information to the server. This data is sent to the server in real time or at specified intervals. The API used for transmission is encrypted as a security measure.
[0844] Step 5:
[0845] The server stores the received voice, gyro sensor, and emotion information in a database, which can then be compared with past data for analysis to detect anomalies.
[0846] Step 6:
[0847] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[0848] Step 7:
[0849] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[0850] Step 8:
[0851] The server analyzes the emotional information and evaluates whether the user is under high stress or fear based on the data provided by the emotion engine.
[0852] Step 9:
[0853] The server uses the results of voice, gyro sensor, and emotion analysis to determine whether an anomaly has been detected, and if so, evaluates whether the anomaly is minor or major.
[0854] Step 10:
[0855] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[0856] Step 11:
[0857] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[0858] Step 12:
[0859] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[0860] Step 13:
[0861] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[0862] Example 2
[0863] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0864] Ensuring the safety of elderly people and children in modern society is a critical issue, requiring rapid and accurate responses, especially in emergencies. Conventional anomaly detection systems respond based solely on voice and motion information, and do not take into account changes in emotions. Such systems have difficulty determining the level of urgency, and there is a risk of delays in issuing appropriate alerts or notifications. Therefore, in order to more reliably protect the personal safety of elderly people and children, there is a need to develop a system that can also comprehensively analyze emotional information.
[0865] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0866] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information to an emotion engine to analyze the user's emotion, means for transmitting the analyzed emotion information to the server, means for transmitting the voice information and gyro sensor information to the server, means for the server to analyze the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, and means for activating a security buzzer when an abnormality is detected. This makes it possible to comprehensively analyze the user's emotion information, appropriately determine which abnormalities are of high urgency, and respond quickly.
[0867] The "means for acquiring voice information" is a means for collecting voice data using a built-in microphone or other voice input device.
[0868] "Means for acquiring gyro sensor information" refers to a means for using a gyro sensor to measure fluctuations in the user's movements and posture and collect that data.
[0869] The "means for transmitting acquired voice information and gyro sensor information to a server" refers to a means having the function of transmitting collected voice data and gyro data to a server via wireless communication or cable connection.
[0870] "Means for the server to analyze audio information and gyro sensor information to detect abnormalities" refers to means for analyzing audio data and gyro data within the server and executing algorithms or programs that detect abnormalities based on the results.
[0871] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to a means that has the function of sending a notification to the user and their pre-designated relatives when an abnormality is detected.
[0872] "Means for sending a notification to the police when a serious abnormality is detected" refers to means that has the function of sending an emergency notification to public institutions such as the police when a highly urgent abnormality is detected.
[0873] "Means for activating the security alarm when an abnormality is detected" refers to means having the function of sounding the security alarm of the terminal to notify those around when the system detects an abnormality.
[0874] An "emotion engine that recognizes user emotions" is an algorithm or program that analyzes acquired voice data to identify and identify the user's emotional state.
[0875] The "means for transmitting acquired voice information to an emotion engine and analyzing the user's emotions" refers to a means having a function for transmitting collected voice data to an emotion engine and analyzing the emotional state.
[0876] The "means for transmitting analyzed emotion information to a server" is a means having a function for transmitting emotion data analyzed by an emotion engine to a server.
[0877] The system of the present invention collects user voice and gyro sensor information, sends it to a server for analysis, detects abnormalities, and issues appropriate alerts and notifications. It also has the ability to collect and analyze user emotional information to more accurately determine the urgency of the abnormality.
[0878] Hardware and software used
[0879] This system uses the following hardware and software:
[0880] Hardware:
[0881] 1. Device (smartphone): Collects voice information using the built-in microphone and acquires movement data using the built-in gyro sensor.
[0882] 2. Server: Analyzes data, detects anomalies, and manages alerts and notifications.
[0883] software:
[0884] 1. Emotion engine: Software for emotion recognition (e.g., IBM Watson or Google Cloud Speech-to-Text).
[0885] 2. Machine learning model: Anomaly detection model trained using TensorFlow and PyTorch.
[0886] 3. Data transmission and management application: An application that manages data transmission from the terminal to the server and instructions from the server.
[0887] Data Acquisition and Processing
[0888] The device uses a built-in microphone to collect audio information at regular intervals (e.g., every 5 seconds). The collected audio data is temporarily stored in the device. Similarly, the built-in gyro sensor is used to obtain motion data, which is also stored in the device.
[0889] The device then sends the collected voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with the collected gyro data.
[0890] The server uses voice analysis algorithms and machine learning models to analyze the received voice information, gyro data, and emotional information. If an abnormality is detected, the server triggers an appropriate alert depending on the urgency. If the abnormality is minor, the server notifies the user and their relatives, and if it is serious, it notifies the police. If a serious abnormality is detected, the server sends an instruction to the device to sound the security alarm.
[0891] Specific examples
[0892] Example 1: An elderly person suddenly collapses at home
[0893] The device's gyro sensor detects violent movements and sends gyro data to the server as a falling motion. At the same time, voice information calling for "help" is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server in real time. The server analyzes this information using a voice analysis algorithm and machine learning model and determines that it is a serious abnormality. The server then sends a notification to family members and alerts the police. It can also send instructions to the device to activate a security alarm, quickly alerting people in the vicinity of an abnormality.
[0894] Example 2: A child is nearly kidnapped in a park
[0895] The device detects sudden movements and cries for help, while the emotion engine detects strong fear. This information is sent to the server in real time. The server immediately detects any abnormalities and notifies the police and parents as serious problems. The server also sends instructions to the device to sound a security alarm, alerting people in the vicinity. This encourages rapid response and helps ensure the safety of children.
[0896] Prompt Sentence Examples
[0897] Example prompt: "Write an outline of an emergency alert system for people with disabilities. The system has voice recognition, emotion detection, and anomaly detection capabilities, and the ability to notify users of detected anomalies."
[0898] Sample prompt: "Describe an anomaly detection system for when an elderly person suddenly collapses. The system uses voice, gyro sensor, and emotion data to detect the anomaly."
[0899] In this way, the system of the present invention makes it possible to quickly and effectively protect the safety of elderly people and children by integrating voice information, gyro sensor information, and emotion information.
[0900] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0901] Step 1:
[0902] Data Acquisition
[0903] The device (smartphone) acquires audio information using a built-in microphone. The microphone continuously captures audio and collects data at regular intervals (for example, every 5 seconds). The input is the user's voice, and the output is audio data stored in the device. Similarly, the device's built-in gyro sensor captures the user's movement data in real time. Gyro data is also collected at regular intervals and stored in the device. The input is the user's movement, and the output is gyro sensor data.
[0904] Step 2:
[0905] emotion recognition
[0906] The device sends the acquired voice information in real time to an emotion engine (for example, IBM Watson or Google Cloud Speech-to-Text) to recognize the user's emotions. The input is the voice information stored on the device, and the output is the emotional state (anger, anxiety, sadness, etc.) obtained from the emotion engine. For example, if the user is angry, the emotion engine will recognize "anger" based on the tone and content of the voice. This emotion information is temporarily stored on the device and sent to the server along with gyro data.
[0907] Step 3:
[0908] Sending data to the server
[0909] The device sends voice information, gyro sensor information, and emotion information to the server. The input is the voice data, gyro data, and emotion information stored on the device, and the output is each data sent to the server. This data transmission is done in real time, but depending on the communication environment, it may also be sent at specified time intervals (for example, every minute).
[0910] Step 4:
[0911] Anomaly detection
[0912] The server analyzes the received voice information, gyro sensor information, and emotional information. The input is the voice data, gyro data, and emotional information sent to the server, and the output is the anomaly detection results. The server is equipped with a machine learning model (for example, a model trained using TensorFlow or PyTorch) and uses this to detect anomalies. Specifically, the server uses a voice analysis algorithm to detect whether the voice data contains specific keywords (for example, "help" or screams). It also analyzes patterns in the gyro data to determine whether the user has suddenly fallen. Furthermore, it evaluates whether the user is experiencing strong stress or fear based on the emotional information provided by the emotion engine.
[0913] Step 5:
[0914] Sending alerts
[0915] If an anomaly is detected, the server triggers an alert depending on its urgency. The input is the anomaly detection result, and the output is the alert to be sent. If the anomaly is minor (for example, a small noise or a slight fall), the server will send an SMS or app notification to the user and their relatives. If the anomaly is serious (for example, if a cry for help or strong emotional expression is detected), the server will immediately notify the police. The notification also includes location information.
[0916] Step 6:
[0917] Activating the security alarm
[0918] When a serious abnormality is detected, the server sends a command to the device to activate the security buzzer. The input is the serious abnormality detection result, and the output is the sound of the security buzzer. The device receives this command and sounds the security buzzer using its built-in speaker. This alerts people in the vicinity to the abnormality and encourages them to take early action.
[0919] (Application example 2)
[0920] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0921] Conventional security systems rely solely on audio and motion information to detect abnormalities, resulting in low accuracy in detecting abnormalities and often making it difficult to determine whether an abnormality has occurred because they do not take into account the emotional state of the user. Furthermore, even if an abnormality is detected, prompt action may not be possible, posing a challenge to ensuring user safety.
[0922] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0923] The smart glasses include a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, a means for activating a security buzzer when an abnormality is detected, a means including an emotion engine that recognizes the user's emotions, a means for transmitting emotion information recognized by the emotion engine to a server, and a means for supporting the safety of users using the smart glasses. This improves the accuracy of abnormality detection and enables prompt and appropriate responses.
[0924] A "means for acquiring voice information" is a device or mechanism that detects a user's voice and records and collects it as digital data.
[0925] "Means for acquiring gyro sensor information" refers to devices or mechanisms that measure changes in a user's movements and posture, and record and collect them as digital data.
[0926] The "means for transmitting to a server" refers to a device or method for transmitting the acquired voice information and gyro sensor information to a server via a communication means such as wireless communication or the Internet.
[0927] "Means for analyzing and detecting anomalies" refers to a processing system that analyzes the data received by the server using algorithms, machine learning models, etc., to identify anomalies.
[0928] An "alert sending means" is a device or method for sending a warning or caution notification to a designated recipient when an abnormality is detected.
[0929] "Means for sending notifications" refers to devices or methods for promptly contacting the user's relatives, police, or other relevant authorities in the event of a serious abnormality.
[0930] "Means for activating a security alarm" refers to the mechanism or control method for activating a physical audio alarm device (security alarm) when an abnormality is detected.
[0931] An "emotion engine" is an algorithm or software that analyzes acquired voice information and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.).
[0932] "Smart glasses" are high-performance glasses that can acquire audio information and gyro sensor information, and can also recognize the user's emotions using an emotion engine, displaying visual information and communicating.
[0933] In this invention, we will build a system in which smart glasses and other applicable devices acquire voice information and gyro sensor information and send it to a server. This allows for quick and appropriate response when an abnormality occurs. Specific implementation methods are shown below.
[0934] First, we use smart glasses as the hardware. The smart glasses have built-in microphones and gyro sensors, which enable them to collect voice and motion information. They also have wireless communication capabilities such as Wi-Fi and Bluetooth, which allow them to send collected data to a server. Second, we use Google Cloud's Speech-to-Text API and Emotion API as emotion engines for emotion recognition.
[0935] The data received by the server is:
[0936] 1. Audio information: The user's voice is picked up by a microphone and recorded as digital data.
[0937] 2. Gyro sensor information: Detects user movements and collects their movement data.
[0938] 3. Emotion information: Emotion data analyzed by the emotion engine.
[0939] The server analyzes the received data and detects anomalies using machine learning models (for example, TensorFlow or PyTorch). Based on the analysis results, it determines the urgency of the anomaly and sends an appropriate alert. Specifically, it notifies the user and their relatives using Twilio's SMS API or Push Notification API.
[0940] If the abnormality is determined to be serious, the server sends a notification to the police and instructs the device to activate a security alarm, which activates a physical audio alarm and alerts people in the vicinity to the abnormality.
[0941] As a concrete example, consider the case where an elderly person suddenly falls while out. In this case, the gyro sensor in the smart glasses detects the fall, and the microphone picks up the cry of "Ouch!" The emotion engine distinguishes between strong pain and fear, and sends this data to the server. The server determines this to be a serious abnormality, sends an emergency call to relatives, and activates the security alarm.
[0942] Prompt Sentence Examples
[0943] "Please build a system that collects voice and gyro sensor information in real time to detect when the user falls or has a sudden change in behavior, and sends the data along with the user's emotional state to a server for analysis and alerts when an abnormality occurs."
[0944] In this way, the invention provides a system that uses smart glasses to collect voice information and gyro sensor information, and also integrates and analyzes emotional information, thereby improving the accuracy of anomaly detection and enabling prompt and appropriate responses.
[0945] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0946] Step 1: Acquire audio and gyro sensor information
[0947] The device (smart glasses) collects the user's voice information using a built-in microphone. At the same time, it acquires the user's movement data using a gyro sensor. The acquired voice information and gyro sensor information are temporarily stored in the device. This input data is the raw data for monitoring the user's voice and movement in real time.
[0948] Step 2: Emotion Recognition
[0949] The device sends the collected voice information to an emotion engine in real time to analyze the user's emotions. The emotion engine (for example, Google Cloud's Speech-to-Text API or Emotion API) analyzes the voice data and identifies the emotional state (anger, anxiety, sadness, etc.). It receives voice information as input data and outputs emotional information. This emotional information is also temporarily stored on the device.
[0950] Step 3: Send data
[0951] The device transmits the acquired voice information, gyro sensor information, and emotion information to the server via Wi-Fi or Bluetooth. This data is sent to the server as input data for analysis.
[0952] Step 4: Anomaly detection
[0953] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models (TensorFlow and PyTorch) to detect whether specific keywords (such as "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the data patterns from the gyro sensor. It also evaluates the user's level of stress based on the emotion information provided by the emotion engine. The analysis results indicate whether there is an abnormality and its urgency.
[0954] Step 5: Sending an alert
[0955] If an abnormality is detected, an alert is sent according to the urgency. If the abnormality is minor, the server sends an alert to the user and their relatives. SMS API and Push Notification API are used for sending the alert. If a serious abnormality occurs, the server also sends a notification to the police. The input data is the analysis result, and the output data is the alert notification.
[0956] Step 6: Activate the security alarm
[0957] If a serious abnormality is detected, the server sends an instruction to the device to activate the security alarm. This activates the security alarm in the smart glasses and alerts people nearby. The input data is the instruction from the server, and the output data is the activation of the security alarm.
[0958] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0959] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0960] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0961] [Fourth embodiment]
[0962] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0963] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0964] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0965] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0966] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0967] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0968] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0969] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0970] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0971] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0972] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0973] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0974] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0975] The present invention is a system that includes a means for acquiring audio information, a means for acquiring gyro sensor information, a means for transmitting the acquired audio information and gyro sensor information to a server, a means for the server to analyze the audio information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected.
[0976] Overview of program processing
[0977] 1. Data Acquisition
[0978] The device (smartphone) acquires voice information using a built-in microphone. This is done using a dedicated application on the device. The voice information is collected at regular intervals and saved as data. The device also acquires information from its built-in gyro sensor in real time to monitor the user's movements and tilt. This gyro sensor information is also saved at regular intervals.
[0979] 2. Sending data to the server
[0980] The voice and gyro sensor information collected by the device is sent to a server in real time or at regular intervals, and the server receives the data and stores it in a database.
[0981] 3. Anomaly detection
[0982] The server analyzes audio and gyro sensor information to detect abnormalities. This analysis is performed using a machine learning model. Audio analysis detects specific keywords (such as screams or "help me"). Gyro sensor analysis also detects sudden movements and unnatural fluctuations (such as falling or sudden changes in acceleration).
[0983] 4. Sending alerts
[0984] If an abnormality is detected, the server evaluates it, and if it is a minor abnormality, it sends an alert to the user and their relatives. This alert is sent via email or app notification. If a serious abnormality is detected, the server also notifies the police, allowing for a prompt emergency response.
[0985] 5. Activating the security alarm
[0986] If an abnormality is detected and is deemed to be particularly serious, the server sends a command to the device to activate the security alarm, thereby notifying people around the user of the abnormality and alerting them to the emergency.
[0987] Specific examples
[0988] Below is a concrete example of how the system of the present invention actually works.
[0989] Consider the case where an elderly person suddenly collapses. At this time, the gyro sensor on the device detects the sudden movement and sends the gyro sensor information to the server as a falling motion. The server analyzes the received gyro information and detects an abnormality such as a fall. If the server determines that this abnormality is serious, it notifies the family and also reports the incident to the police. The server also sends instructions to the device to activate the security alarm. This alerts people in the vicinity to the abnormality, enabling a swift response.
[0990] The system also simulates a child being abducted. In this case, the gyro sensor detects any sudden movements or changes in behavior, and the child's cries for help are captured as audio information. This information is sent to the server, which analyzes it to detect any abnormalities. The server determines this to be a very serious anomaly and immediately notifies the police. At the same time, a notification is sent to the parents, and a security alarm is activated. This is expected to allow for early response.
[0991] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[0992] The processing flow will be explained below.
[0993] Step 1:
[0994] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[0995] Step 2:
[0996] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[0997] Step 3:
[0998] The device sends the collected voice and gyro sensor information to a server. In particular, this data is sent to the server in real time or at specified intervals via internet communication. The API used for transmission is encrypted as a security measure.
[0999] Step 4:
[1000] The server stores the received audio and gyro sensor information in a database, which allows analysis to detect anomalies by comparing it with past data.
[1001] Step 5:
[1002] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[1003] Step 6:
[1004] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[1005] Step 7:
[1006] The server determines whether an abnormality is detected based on the results of voice analysis and gyro sensor analysis, and if so, evaluates whether the abnormality is minor or serious.
[1007] Step 8:
[1008] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[1009] Step 9:
[1010] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[1011] Step 10:
[1012] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[1013] Step 11:
[1014] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[1015] Example 1
[1016] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1017] In modern society, when vulnerable people such as the elderly and children face an emergency, there is a need to respond quickly and appropriately, but there are still few systems that can achieve this, and existing technologies are insufficient. In particular, there is a lack of means to detect serious abnormalities such as falls and kidnappings early and to take prompt and appropriate action.
[1018] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1019] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information and gyro sensor information to the server, means for analyzing the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, means for activating a security buzzer when an abnormality is detected, means for sending the voice information and gyro sensor information acquired by the terminal to the server via Wi-Fi or a mobile data network, and means for collecting and storing various sensor information at time intervals. This makes it possible to quickly collect and analyze voice information and gyro sensor information and take appropriate action when elderly people, children, or others face an emergency.
[1020] "Means for acquiring audio information" refers to the means by which a terminal uses a built-in microphone or other audio recording device to collect surrounding audio as digital data.
[1021] "Means for acquiring gyro sensor information" refers to the means by which a device uses its built-in gyro sensor to collect data on the user's movements and tilt in real time.
[1022] "Means for transmitting acquired voice information and gyro sensor information to a server" refers to means for transmitting voice information and gyro sensor information collected by the terminal to a server via a communication means such as Wi-Fi or a mobile data network.
[1023] "Means for the server to analyze audio information and gyro sensor information to detect anomalies" refers to means for the server to use machine learning models or other analytical technologies to analyze received audio information and gyro sensor information and detect anomalies.
[1024] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to the means for informing the user and their relatives of an abnormality by sending an email or app notification when the server detects an abnormality.
[1025] The "means for sending a notification to the police when a serious abnormality is detected" is a means for the server to quickly notify the police when it detects a serious abnormality.
[1026] The "means for activating the security buzzer when an abnormality is detected" is a means for the server to send an instruction to the terminal to activate the security buzzer of the terminal when an abnormality is detected.
[1027] "Means for transmitting voice information and gyro sensor information acquired by a terminal to a server via Wi-Fi or a mobile data network" refers to means for transmitting voice information and gyro sensor information collected by a terminal to a server using wireless communication technology.
[1028] "Means for collecting and storing various sensor information at regular intervals" refers to means by which the terminal automatically collects and stores audio information and gyro sensor information at regular intervals.
[1029] This invention is a system for responding quickly and appropriately to emergencies faced by vulnerable people such as the elderly and children. This system acquires voice and gyro sensor information, transmits it to a server for analysis, and activates an alert or a security buzzer when an abnormality is detected.
[1030] Data Acquisition
[1031] The device (smartphone) acquires audio information using a built-in microphone. This is done using a dedicated application on the smartphone. The audio information is collected at regular intervals (for example, every 5 seconds) and saved as data. The device also uses a built-in gyro sensor to monitor the user's movements and tilt in real time. The gyro sensor information is also saved at regular intervals. In this way, the device can acquire audio data and gyro sensor data simultaneously.
[1032] Sending data to the server
[1033] The voice and gyro sensor information collected by the device is sent to a server via Wi-Fi or mobile data network in real time or at regular intervals. The server receives the sent data and stores it in a database. During this process, communication between the device and the server is encrypted to ensure security.
[1034] Anomaly detection
[1035] The server analyzes the received voice information and gyro sensor information. This analysis uses generative AI models and machine learning models. Specifically, natural language processing technology is used to analyze the voice information to detect specific keywords (e.g., "help" or "danger"). Additionally, analysis of the gyro sensor information detects sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration). If an abnormality is detected as a result of the analysis, the system proceeds to the next stage.
[1036] Sending alerts
[1037] If the server detects an abnormality, it evaluates the urgency of that abnormality. If a minor abnormality is detected (for example, if only audio containing specific keywords is detected), the server sends an alert to the user and their relatives. This alert is sent via email or a dedicated application. If a serious abnormality is detected (for example, if audio information and abnormal behavior are detected simultaneously), the server also notifies the police. This allows for a swift emergency response.
[1038] Activating the security alarm
[1039] If the server detects a serious abnormality, it will send a command to the device to activate the security alarm. Activating the security alarm will alert people around the user to the abnormality and make it clear that it is an emergency. This function will enable people around the device to take early action.
[1040] Specific examples
[1041] Falls in the elderly
[1042] If an elderly person suddenly falls, the device's gyro sensor detects the sudden movement and sends that information to the server. The server analyzes the received gyro sensor information and detects the abnormality as a fall. If the server determines that the abnormality is serious, it notifies the family and also reports the incident to the police. The server then sends an instruction to activate the device's security alarm, which then emits a high-volume alarm to alert those around it to the abnormality.
[1043] Child abduction
[1044] In the event of a child abduction, the device's gyro sensor will detect any sudden movements or changes in behavior, and the child's cries for help will be picked up as audio information. This information is sent to a server, which analyzes it to detect any abnormalities. The server will determine this to be a very serious anomaly and immediately notify the police. At the same time, a notification will be sent to the parents, and a security alarm will be activated. This is expected to allow for early response.
[1045] Prompt Sentence Examples
[1046] "Please explain a system that can acquire voice and gyro sensor information and detect abnormalities when an elderly person suddenly collapses."
[1047] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children and students by using voice information and gyro sensor information.
[1048] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1049] Step 1: Get the data
[1050] The device uses a built-in microphone to acquire audio information. This is done using a dedicated application, and audio is collected at regular intervals (e.g., every 5 seconds). The input is ambient audio, and digitized audio data is obtained as output. The device also uses a built-in gyro sensor to monitor the user's movement and tilt. This gyro sensor information is also collected at regular intervals, and acceleration data for the x-, y-, and z-axes is obtained as output.
[1051] Step 2: Send data to the server
[1052] The device sends the collected voice information and gyro sensor information to the server via Wi-Fi or mobile data network. The input is the voice data and gyro sensor data collected by the device, and this data is sent to the server as packets at regular time intervals (e.g., 5 seconds). The output is the data packets sent to the server.
[1053] Step 3: Detect anomalies
[1054] The server analyzes the received voice information and gyro sensor information. Generative AI models and machine learning models are used for this analysis. Based on the voice data input, natural language processing techniques are used to detect specific keywords (e.g., "help," "danger"), and based on the gyro sensor data input, sudden movements and unnatural fluctuations (e.g., falling or sudden changes in acceleration) are detected. The analysis results are obtained as output, and a flag is raised if an abnormality is detected.
[1055] Step 4: Sending an alert
[1056] If an anomaly is detected, the server evaluates the anomaly and sends an alert to the user and their relatives. The input is the analysis result, and if the anomaly flag is true, an action is triggered. If the anomaly is minor, the output is an email or app notification sent to the user and their relatives. If a serious anomaly is detected, a notification is also sent to the police.
[1057] Step 5: Activate the security alarm
[1058] If a serious abnormality is detected, the server sends an instruction to the terminal to activate the security buzzer. The input is the abnormality evaluation result, and if the serious abnormality flag is true, an instruction to activate the security buzzer is sent as output to the terminal. Upon receiving this instruction, the terminal activates the security buzzer and emits a high-volume alarm to alert those in the vicinity of the abnormality.
[1059] (Application example 1)
[1060] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1061] In recent years, ensuring the safety of vulnerable people, including the elderly and children, has become an important issue. However, current security systems often lack the ability to recognize abnormalities in real time or respond quickly. Furthermore, few systems that detect abnormalities using both audio and motion information efficiently integrate these data, resulting in problems such as false detection and delayed response. The present invention aims to solve these problems by providing a system that can quickly and accurately detect and respond to abnormalities, especially when the elderly or children are in danger.
[1062] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1063] In this invention, the server includes means for analyzing voice information and gyro sensor information to detect abnormalities, means for analyzing specific keywords from the voice information and detecting sudden changes in movement from the gyro sensor information, and means for sending an alert and activating a security buzzer if necessary if an abnormality is detected based on the analysis results. This makes it possible to comprehensively monitor the user's voice and movement, respond quickly in the event of an abnormality, and notify the user and their relatives or those around them of an emergency by activating the security buzzer.
[1064] "Audio information" is data that describes the user's environmental sounds and speech content.
[1065] A "gyro sensor" is a device used to detect the user's movement and tilt.
[1066] The "server" is a computer system that analyzes collected audio information and gyro sensor information, and detects and notifies users of abnormalities.
[1067] An "alert" is a notification sent to the user, their relatives, and, if necessary, the police or other relevant parties when an abnormality is detected.
[1068] A "security buzzer" is a device that emits a sound when an abnormality is detected to alert those around to danger.
[1069] "Analysis" is a process for detecting the presence or absence of abnormalities based on collected audio information and gyro sensor information.
[1070] The "specific keywords" are words and phrases that mainly indicate an emergency in the audio information that indicates an abnormality.
[1071] "Sudden changes in movement" refers to sudden body movements that deviate from normal movement.
[1072] The present invention provides a system that uses voice information and gyro sensor information to monitor the safety of a user and responds quickly when an abnormality is detected. Specific embodiments for carrying out the present invention will be described below.
[1073] System Configuration
[1074] Hardware:
[1075] Smartphone: Equipped with a built-in microphone and gyro sensor, it acquires voice and gyro sensor information.
[1076] Server: Analyzes audio and gyro sensor information to detect abnormalities, send notifications, activate the security alarm, etc.
[1077] software:
[1078] Dedicated application: Equipped with functions to collect voice information and gyro sensor information and send it to a server.
[1079] Cloud server: Uses AWS, GCP, Azure, etc. to analyze and manage data.
[1080] Machine learning model: Using Python or similar software, an algorithm is implemented to analyze audio information and gyro sensor information and detect anomalies.
[1081] Speech recognition technology: Using technologies such as IBM Watson and Google Speech-to-Text, voice data is converted into text and specific keywords are detected.
[1082] Program Processing Overview
[1083] Acquiring audio information
[1084] The user's smartphone uses a built-in microphone to capture voice information, which is then sent to a server at regular intervals via a dedicated application.
[1085] Get gyro sensor information
[1086] The smartphone's built-in gyro sensor is used to monitor the user's movements and tilt in real time, and the application periodically sends the data to a server.
[1087] Data analysis on the server
[1088] The server analyzes the received voice and gyro sensor information. The voice information is converted into text using speech recognition technology to detect whether it contains specific keywords (such as "help" or screams). The gyro sensor information is analyzed using a machine learning model to detect sudden changes in movement (e.g., falling).
[1089] Sending alerts
[1090] If an abnormality is detected, the server evaluates its severity, and if it is a minor abnormality, it sends an app notification to the user and their relatives. If it is a serious abnormality, it also notifies the police.
[1091] Activating the security alarm
[1092] If a serious abnormality is detected, a security alarm will be activated on the smartphone, alerting people nearby.
[1093] Specific examples
[1094] Falls in the elderly
[1095] If an elderly person suddenly falls, the smartphone's gyro sensor will detect the sudden movement and send the information to the server. The server will then detect the fall, notify the family, and notify the police. It will also activate a security alarm to alert people nearby of the emergency.
[1096] Child abduction prevention
[1097] In the case of a child being kidnapped, the smartphone's gyro sensor will detect any sudden changes in movement and capture any voices crying for help. This information will be sent to a server, which will then analyze it and determine that the child has been kidnapped. The server will then immediately notify the police, and a notification will be sent to the parents, who will then activate the security alarm. This is expected to allow for early response.
[1098] Prompt Sentence Examples
[1099] I'd like to see some code examples for analyzing data from an app to detect anomalies in user voice and behavior. The voice data needs to detect specific keywords, such as "help" or "scream." Also, the gyro sensor information needs to detect sudden changes in behavior. I'll be using Python.
[1100] As described above, the system of the present invention can quickly and effectively protect the safety of elderly people and children by using voice information and gyro sensor information.
[1101] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1102] Step 1:
[1103] Acquiring audio information
[1104] The user acquires audio information using a smartphone. The smartphone's built-in microphone collects the user's ambient sounds and speech, and saves the audio data at regular intervals through a dedicated application. The input is the ambient sounds and the user's speech, and the output is the saved audio data.
[1105] Step 2:
[1106] Get gyro sensor information
[1107] The device uses a built-in gyro sensor to monitor the user's movements in real time. The gyro sensor captures acceleration information along the x-, y-, and z-axes, and periodically saves this as data through a dedicated application. The input is the user's movements, and the output is the captured gyro sensor information.
[1108] Step 3:
[1109] Voice information and gyro sensor information sent to the server
[1110] The smartphone transmits the stored voice and gyro sensor information to the server at regular intervals using a secure communication protocol (such as HTTPS). The input is the voice data and gyro sensor information, and the output is the data transmitted to the server.
[1111] Step 4:
[1112] Saving data to the server
[1113] The server stores the received audio information and gyro sensor information in a database. The input is the transmitted data, and the output is the data stored in the database.
[1114] Step 5:
[1115] Analysis of audio information
[1116] The server uses speech recognition technology to convert the stored voice information into text, and analyzes this text data to see if it contains specific keywords (such as "help" or a cry). The input is the stored voice information, and the output is the text data and the analysis results.
[1117] Step 6:
[1118] Analysis of gyro sensor information
[1119] The server analyzes the gyro sensor information using a machine learning model, detects sudden changes in movement (e.g., falling or sudden acceleration), and evaluates the user's situation. The input is the stored gyro sensor information, and the output is the analysis result.
[1120] Step 7:
[1121] Anomaly detection
[1122] The server detects abnormalities based on the results of analyzing voice information and gyro sensor information. It evaluates the degree of abnormality based on the detection results and decides how to respond. The input is the results of voice analysis and gyro analysis, and the output is the presence or absence of an abnormality and its degree.
[1123] Step 8:
[1124] Sending alerts
[1125] The server sends an alert to the user and their relatives depending on the detected anomaly. If the anomaly is minor, it sends an app notification, and if it is severe, it also notifies the police. The input is the degree of anomaly, and the output is the alert sent.
[1126] Step 9:
[1127] Activating the security alarm
[1128] If the server determines that a serious abnormality exists, it sends a command to the smartphone to activate the security buzzer. This notifies people around the smartphone of the emergency. The input is the degree of abnormality, and the output is the activation of the security buzzer.
[1129] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1130] The present invention is a system that includes a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, and a means for activating a security buzzer when an abnormality is detected, and further includes an emotion engine that recognizes the user's emotions.
[1131] Overview of program processing
[1132] 1. Data Acquisition
[1133] The device (smartphone) acquires voice information using a built-in microphone. The voice information is collected at regular intervals and temporarily stored within the device. Motion data is also acquired using a gyro sensor and is also stored within the device.
[1134] 2. Emotion recognition
[1135] The device transmits the acquired voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with other sensor data.
[1136] 3. Sending data to the server
[1137] The device sends collected voice information, gyro sensor information, and emotion information to a server, which sends this data in real time or at specified intervals.
[1138] 4. Anomaly Detection
[1139] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models to detect whether specific keywords (e.g., "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data. It also evaluates whether the user is under high stress based on the emotion information provided by the emotion engine.
[1140] 5. Sending alerts
[1141] If an anomaly is detected, the server will trigger an appropriate alert depending on the severity: if it is a minor anomaly, an alert will be sent to the user and their relatives; if it is a serious anomaly, a notification will also be sent to the police.
[1142] 6. Activating the security alarm
[1143] If a serious abnormality is detected, the server sends a command to the device to activate the burglar alarm, which then automatically sounds the alarm to alert people in the vicinity.
[1144] Specific examples
[1145] For example, let's simulate an elderly person suddenly collapsing at home. The device's gyro sensor detects the sudden movement and sends gyro data to the server as a falling motion. At the same time, a voice calling for help is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server, and the analysis results indicate a serious abnormality. The server notifies the family and also alerts the police. It also sends instructions to the device to activate the security alarm. This allows people in the vicinity to be quickly notified of the abnormality.
[1146] The system also simulates a situation where a child is about to be kidnapped in a park. In this case, the device detects sudden movements and the voice calling for help, while the emotion engine detects strong fear. This information is sent to the server, which determines it as a serious abnormality. The server then immediately notifies the police and parents and instructs the device to activate the security alarm, enabling early action to be taken and helping to ensure the child's safety.
[1147] In this way, the system of the present invention can quickly and effectively protect the safety of the elderly and children by integrating voice information, gyro sensor information, and emotional information.
[1148] The processing flow will be explained below.
[1149] Step 1:
[1150] The device acquires audio information. Specifically, it uses a built-in microphone to record surrounding audio at regular intervals. This audio data is temporarily stored in the device.
[1151] Step 2:
[1152] The device acquires gyro sensor information. Using the device's built-in gyro sensor, it collects motion data such as acceleration and tilt in real time. This data is also temporarily stored on the device.
[1153] Step 3:
[1154] The device sends the collected voice information to the emotion engine, which analyzes the voice data in real time and recognizes the user's emotional state (e.g., anger, anxiety, sadness, etc.). The recognized emotion information is temporarily stored in the device along with other sensor data.
[1155] Step 4:
[1156] The device transmits voice information, gyro sensor information, and emotion information to the server. This data is sent to the server in real time or at specified intervals. The API used for transmission is encrypted as a security measure.
[1157] Step 5:
[1158] The server stores the received voice, gyro sensor, and emotion information in a database, which can then be compared with past data for analysis to detect anomalies.
[1159] Step 6:
[1160] The server analyzes the audio information, using speech analysis algorithms and machine learning models to detect whether it contains certain keywords (e.g., "help" or screams).
[1161] Step 7:
[1162] The server analyzes the gyro sensor information and uses an algorithm to determine whether a fall or a sudden change in movement has occurred based on the fluctuation pattern of the gyro data.
[1163] Step 8:
[1164] The server analyzes the emotional information and evaluates whether the user is under high stress or fear based on the data provided by the emotion engine.
[1165] Step 9:
[1166] The server uses the results of voice, gyro sensor, and emotion analysis to determine whether an anomaly has been detected, and if so, evaluates whether the anomaly is minor or major.
[1167] Step 10:
[1168] If the server detects any minor abnormalities, it will send an alert to the user and their relatives via email or app notification.
[1169] Step 11:
[1170] If a serious anomaly is detected, the server will also notify the police, who will be notified promptly via emergency notification systems or APIs.
[1171] Step 12:
[1172] The server sends a command to the device to activate the security alarm, which causes the device to automatically sound the alarm and alert people in the vicinity to an abnormality.
[1173] Step 13:
[1174] The user, their relatives, or the police will be notified and will quickly rush to the scene to deal with the emergency. The alarm will also sound, making it easier to get help from those around.
[1175] Example 2
[1176] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1177] Ensuring the safety of elderly people and children in modern society is a critical issue, requiring rapid and accurate responses, especially in emergencies. Conventional anomaly detection systems respond based solely on voice and motion information, and do not take into account changes in emotions. Such systems have difficulty determining the level of urgency, and there is a risk of delays in issuing appropriate alerts or notifications. Therefore, in order to more reliably protect the personal safety of elderly people and children, there is a need to develop a system that can also comprehensively analyze emotional information.
[1178] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1179] In this invention, the server includes means for acquiring voice information, means for acquiring gyro sensor information, means for transmitting the acquired voice information to an emotion engine to analyze the user's emotion, means for transmitting the analyzed emotion information to the server, means for transmitting the voice information and gyro sensor information to the server, means for the server to analyze the voice information and gyro sensor information to detect abnormalities, means for sending an alert to the user and their relatives when an abnormality is detected, means for sending a notification to the police when a serious abnormality is detected, and means for activating a security buzzer when an abnormality is detected. This makes it possible to comprehensively analyze the user's emotion information, appropriately determine which abnormalities are of high urgency, and respond quickly.
[1180] The "means for acquiring voice information" is a means for collecting voice data using a built-in microphone or other voice input device.
[1181] "Means for acquiring gyro sensor information" refers to a means for using a gyro sensor to measure fluctuations in the user's movements and posture and collect that data.
[1182] The "means for transmitting acquired voice information and gyro sensor information to a server" refers to a means having the function of transmitting collected voice data and gyro data to a server via wireless communication or cable connection.
[1183] "Means for the server to analyze audio information and gyro sensor information to detect abnormalities" refers to means for analyzing audio data and gyro data within the server and executing algorithms or programs that detect abnormalities based on the results.
[1184] "Means for sending an alert to the user and their relatives when an abnormality is detected" refers to a means that has the function of sending a notification to the user and their pre-designated relatives when an abnormality is detected.
[1185] "Means for sending a notification to the police when a serious abnormality is detected" refers to means that has the function of sending an emergency notification to public institutions such as the police when a highly urgent abnormality is detected.
[1186] "Means for activating the security alarm when an abnormality is detected" refers to means having the function of sounding the security alarm of the terminal to notify those around when the system detects an abnormality.
[1187] An "emotion engine that recognizes user emotions" is an algorithm or program that analyzes acquired voice data to identify and identify the user's emotional state.
[1188] The "means for transmitting acquired voice information to an emotion engine and analyzing the user's emotions" refers to a means having a function for transmitting collected voice data to an emotion engine and analyzing the emotional state.
[1189] The "means for transmitting analyzed emotion information to a server" is a means having a function for transmitting emotion data analyzed by an emotion engine to a server.
[1190] The system of the present invention collects user voice and gyro sensor information, sends it to a server for analysis, detects abnormalities, and issues appropriate alerts and notifications. It also has the ability to collect and analyze user emotional information to more accurately determine the urgency of the abnormality.
[1191] Hardware and software used
[1192] This system uses the following hardware and software:
[1193] Hardware:
[1194] 1. Device (smartphone): Collects voice information using the built-in microphone and acquires movement data using the built-in gyro sensor.
[1195] 2. Server: Analyzes data, detects anomalies, and manages alerts and notifications.
[1196] software:
[1197] 1. Emotion engine: Software for emotion recognition (e.g., IBM Watson or Google Cloud Speech-to-Text).
[1198] 2. Machine learning model: Anomaly detection model trained using TensorFlow and PyTorch.
[1199] 3. Data transmission and management application: An application that manages data transmission from the terminal to the server and instructions from the server.
[1200] Data Acquisition and Processing
[1201] The device uses a built-in microphone to collect audio information at regular intervals (e.g., every 5 seconds). The collected audio data is temporarily stored in the device. Similarly, the built-in gyro sensor is used to obtain motion data, which is also stored in the device.
[1202] The device then sends the collected voice information to the emotion engine in real time to recognize the user's emotions. The emotion engine analyzes the voice data and identifies the user's emotional state (anger, anxiety, sadness, etc.). The recognized emotion information is sent to the server along with the collected gyro data.
[1203] The server uses voice analysis algorithms and machine learning models to analyze the received voice information, gyro data, and emotional information. If an abnormality is detected, the server triggers an appropriate alert depending on the urgency. If the abnormality is minor, the server notifies the user and their relatives, and if it is serious, it notifies the police. If a serious abnormality is detected, the server sends an instruction to the device to sound the security alarm.
[1204] Specific examples
[1205] Example 1: An elderly person suddenly collapses at home
[1206] The device's gyro sensor detects violent movements and sends gyro data to the server as a falling motion. At the same time, voice information calling for "help" is detected, and the emotion engine detects strong fear or anxiety. This data is sent to the server in real time. The server analyzes this information using a voice analysis algorithm and machine learning model and determines that it is a serious abnormality. The server then sends a notification to family members and alerts the police. It can also send instructions to the device to activate a security alarm, quickly alerting people in the vicinity of an abnormality.
[1207] Example 2: A child is nearly kidnapped in a park
[1208] The device detects sudden movements and cries for help, while the emotion engine detects strong fear. This information is sent to the server in real time. The server immediately detects any abnormalities and notifies the police and parents as serious problems. The server also sends instructions to the device to sound a security alarm, alerting people in the vicinity. This encourages rapid response and helps ensure the safety of children.
[1209] Prompt Sentence Examples
[1210] Example prompt: "Write an outline of an emergency alert system for people with disabilities. The system has voice recognition, emotion detection, and anomaly detection capabilities, and the ability to notify users of detected anomalies."
[1211] Sample prompt: "Describe an anomaly detection system for when an elderly person suddenly collapses. The system uses voice, gyro sensor, and emotion data to detect the anomaly."
[1212] In this way, the system of the present invention makes it possible to quickly and effectively protect the safety of elderly people and children by integrating voice information, gyro sensor information, and emotion information.
[1213] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1214] Step 1:
[1215] Data Acquisition
[1216] The device (smartphone) acquires audio information using a built-in microphone. The microphone continuously captures audio and collects data at regular intervals (for example, every 5 seconds). The input is the user's voice, and the output is audio data stored in the device. Similarly, the device's built-in gyro sensor captures the user's movement data in real time. Gyro data is also collected at regular intervals and stored in the device. The input is the user's movement, and the output is gyro sensor data.
[1217] Step 2:
[1218] emotion recognition
[1219] The device sends the acquired voice information in real time to an emotion engine (for example, IBM Watson or Google Cloud Speech-to-Text) to recognize the user's emotions. The input is the voice information stored on the device, and the output is the emotional state (anger, anxiety, sadness, etc.) obtained from the emotion engine. For example, if the user is angry, the emotion engine will recognize "anger" based on the tone and content of the voice. This emotion information is temporarily stored on the device and sent to the server along with gyro data.
[1220] Step 3:
[1221] Sending data to the server
[1222] The device sends voice information, gyro sensor information, and emotion information to the server. The input is the voice data, gyro data, and emotion information stored on the device, and the output is each data sent to the server. This data transmission is done in real time, but depending on the communication environment, it may also be sent at specified time intervals (for example, every minute).
[1223] Step 4:
[1224] Anomaly detection
[1225] The server analyzes the received voice information, gyro sensor information, and emotional information. The input is the voice data, gyro data, and emotional information sent to the server, and the output is the anomaly detection results. The server is equipped with a machine learning model (for example, a model trained using TensorFlow or PyTorch) and uses this to detect anomalies. Specifically, the server uses a voice analysis algorithm to detect whether the voice data contains specific keywords (for example, "help" or screams). It also analyzes patterns in the gyro data to determine whether the user has suddenly fallen. Furthermore, it evaluates whether the user is experiencing strong stress or fear based on the emotional information provided by the emotion engine.
[1226] Step 5:
[1227] Sending alerts
[1228] If an anomaly is detected, the server triggers an alert depending on its urgency. The input is the anomaly detection result, and the output is the alert to be sent. If the anomaly is minor (for example, a small noise or a slight fall), the server will send an SMS or app notification to the user and their relatives. If the anomaly is serious (for example, if a cry for help or strong emotional expression is detected), the server will immediately notify the police. The notification also includes location information.
[1229] Step 6:
[1230] Activating the security alarm
[1231] When a serious abnormality is detected, the server sends a command to the device to activate the security buzzer. The input is the serious abnormality detection result, and the output is the sound of the security buzzer. The device receives this command and sounds the security buzzer using its built-in speaker. This alerts people in the vicinity to the abnormality and encourages them to take early action.
[1232] (Application example 2)
[1233] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1234] Conventional security systems rely solely on audio and motion information to detect abnormalities, resulting in low accuracy in detecting abnormalities and often making it difficult to determine whether an abnormality has occurred because they do not take into account the emotional state of the user. Furthermore, even if an abnormality is detected, prompt action may not be possible, posing a challenge to ensuring user safety.
[1235] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1236] The smart glasses include a means for acquiring voice information, a means for acquiring gyro sensor information, a means for transmitting the acquired voice information and gyro sensor information to a server, a means for the server to analyze the voice information and gyro sensor information to detect abnormalities, a means for sending an alert to the user and their relatives when an abnormality is detected, a means for sending a notification to the police when a serious abnormality is detected, a means for activating a security buzzer when an abnormality is detected, a means including an emotion engine that recognizes the user's emotions, a means for transmitting emotion information recognized by the emotion engine to a server, and a means for supporting the safety of users using the smart glasses. This improves the accuracy of abnormality detection and enables prompt and appropriate responses.
[1237] A "means for acquiring voice information" is a device or mechanism that detects a user's voice and records and collects it as digital data.
[1238] "Means for acquiring gyro sensor information" refers to devices or mechanisms that measure changes in a user's movements and posture, and record and collect them as digital data.
[1239] The "means for transmitting to a server" refers to a device or method for transmitting the acquired voice information and gyro sensor information to a server via a communication means such as wireless communication or the Internet.
[1240] "Means for analyzing and detecting anomalies" refers to a processing system that analyzes the data received by the server using algorithms, machine learning models, etc., to identify anomalies.
[1241] An "alert sending means" is a device or method for sending a warning or caution notification to a designated recipient when an abnormality is detected.
[1242] "Means for sending notifications" refers to devices or methods for promptly contacting the user's relatives, police, or other relevant authorities in the event of a serious abnormality.
[1243] "Means for activating a security alarm" refers to the mechanism or control method for activating a physical audio alarm device (security alarm) when an abnormality is detected.
[1244] An "emotion engine" is an algorithm or software that analyzes acquired voice information and identifies the user's emotional state (e.g., anger, anxiety, sadness, etc.).
[1245] "Smart glasses" are high-performance glasses that can acquire audio information and gyro sensor information, and can also recognize the user's emotions using an emotion engine, displaying visual information and communicating.
[1246] In this invention, we will build a system in which smart glasses and other applicable devices acquire voice information and gyro sensor information and send it to a server. This allows for quick and appropriate response when an abnormality occurs. Specific implementation methods are shown below.
[1247] First, we use smart glasses as the hardware. The smart glasses have built-in microphones and gyro sensors, which enable them to collect voice and motion information. They also have wireless communication capabilities such as Wi-Fi and Bluetooth, which allow them to send collected data to a server. Second, we use Google Cloud's Speech-to-Text API and Emotion API as emotion engines for emotion recognition.
[1248] The data received by the server is:
[1249] 1. Audio information: The user's voice is picked up by a microphone and recorded as digital data.
[1250] 2. Gyro sensor information: Detects user movements and collects their movement data.
[1251] 3. Emotion information: Emotion data analyzed by the emotion engine.
[1252] The server analyzes the received data and detects anomalies using machine learning models (for example, TensorFlow or PyTorch). Based on the analysis results, it determines the urgency of the anomaly and sends an appropriate alert. Specifically, it notifies the user and their relatives using Twilio's SMS API or Push Notification API.
[1253] If the abnormality is determined to be serious, the server sends a notification to the police and instructs the device to activate a security alarm, which activates a physical audio alarm and alerts people in the vicinity to the abnormality.
[1254] As a concrete example, consider the case where an elderly person suddenly falls while out. In this case, the gyro sensor in the smart glasses detects the fall, and the microphone picks up the cry of "Ouch!" The emotion engine distinguishes between strong pain and fear, and sends this data to the server. The server determines this to be a serious abnormality, sends an emergency call to relatives, and activates the security alarm.
[1255] Prompt Sentence Examples
[1256] "Please build a system that collects voice and gyro sensor information in real time to detect when the user falls or has a sudden change in behavior, and sends the data along with the user's emotional state to a server for analysis and alerts when an abnormality occurs."
[1257] In this way, the invention provides a system that uses smart glasses to collect voice information and gyro sensor information, and also integrates and analyzes emotional information, thereby improving the accuracy of anomaly detection and enabling prompt and appropriate responses.
[1258] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1259] Step 1: Acquire audio and gyro sensor information
[1260] The device (smart glasses) collects the user's voice information using a built-in microphone. At the same time, it acquires the user's movement data using a gyro sensor. The acquired voice information and gyro sensor information are temporarily stored in the device. This input data is the raw data for monitoring the user's voice and movement in real time.
[1261] Step 2: Emotion Recognition
[1262] The device sends the collected voice information to an emotion engine in real time to analyze the user's emotions. The emotion engine (for example, Google Cloud's Speech-to-Text API or Emotion API) analyzes the voice data and identifies the emotional state (anger, anxiety, sadness, etc.). It receives voice information as input data and outputs emotional information. This emotional information is also temporarily stored on the device.
[1263] Step 3: Send data
[1264] The device transmits the acquired voice information, gyro sensor information, and emotion information to the server via Wi-Fi or Bluetooth. This data is sent to the server as input data for analysis.
[1265] Step 4: Anomaly detection
[1266] The server analyzes the received voice information, gyro sensor information, and emotion information. It uses voice analysis algorithms and machine learning models (TensorFlow and PyTorch) to detect whether specific keywords (such as "help" or screams) are included. It also determines whether a fall or a sudden change in movement has occurred based on the data patterns from the gyro sensor. It also evaluates the user's level of stress based on the emotion information provided by the emotion engine. The analysis results indicate whether there is an abnormality and its urgency.
[1267] Step 5: Sending an alert
[1268] If an abnormality is detected, an alert is sent according to the urgency. If the abnormality is minor, the server sends an alert to the user and their relatives. SMS API and Push Notification API are used for sending the alert. If a serious abnormality occurs, the server also sends a notification to the police. The input data is the analysis result, and the output data is the alert notification.
[1269] Step 6: Activate the security alarm
[1270] If a serious abnormality is detected, the server sends an instruction to the device to activate the security alarm. This activates the security alarm in the smart glasses and alerts people nearby. The input data is the instruction from the server, and the output data is the activation of the security alarm.
[1271] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1272] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1273] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1274] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1275] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1276] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1277] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1278] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1279] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1280] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1281] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1282] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1283] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1284] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1285] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1286] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1287] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1288] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1289] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1290] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1291] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1292] The following is further disclosed regarding the above embodiment.
[1293] (Claim 1)
[1294] A means for acquiring voice information;
[1295] A means for acquiring gyro sensor information;
[1296] means for transmitting the acquired voice information and gyro sensor information to a server;
[1297] A means for the server to analyze the audio information and the gyro sensor information to detect an abnormality;
[1298] A means for sending an alert to the user and their relatives when an anomaly is detected;
[1299] A means of sending a notification to the police when a serious abnormality is detected;
[1300] A means for activating a security alarm when an abnormality is detected;
[1301] A system including:
[1302] (Claim 2)
[1303] 10. The system of claim 1, further comprising means for analyzing the audio information and the gyro sensor information with a machine learning model.
[1304] (Claim 3)
[1305] 2. The system according to claim 1, further comprising means for determining, when an abnormality is detected, a destination of an alert depending on the urgency of the abnormality.
[1306] "Example 1"
[1307] (Claim 1)
[1308] A means for acquiring voice information;
[1309] A means for acquiring gyro sensor information;
[1310] means for transmitting the acquired voice information and gyro sensor information to a server;
[1311] A means for the server to analyze the audio information and the gyro sensor information to detect an abnormality;
[1312] A means for sending an alert to the user and their relatives when an anomaly is detected;
[1313] A means of sending a notification to the police when a serious abnormality is detected;
[1314] A means for activating a security alarm when an abnormality is detected;
[1315] A means for transmitting the voice information and gyro sensor information acquired by the terminal to a server via Wi-Fi or a mobile data network;
[1316] A system that includes a means of collecting and storing various sensor information at time intervals.
[1317] (Claim 2)
[1318] 10. The system of claim 1, further comprising means for analyzing the audio information and the gyro sensor information with a machine learning model.
[1319] (Claim 3)
[1320] 2. The system according to claim 1, further comprising means for determining, when an abnormality is detected, a destination of an alert depending on the urgency of the abnormality.
[1321] "Application Example 1"
[1322] (Claim 1)
[1323] A means for acquiring voice information;
[1324] A means for acquiring gyro sensor information;
[1325] means for transmitting the acquired voice information and gyro sensor information to a server;
[1326] A means for the server to analyze the audio information and the gyro sensor information to detect an abnormality;
[1327] A means for sending an alert to the user and their relatives when an anomaly is detected;
[1328] A means of sending a notification to the police when a serious abnormality is detected;
[1329] A means for activating a security alarm when an abnormality is detected;
[1330] A means for analyzing specific keywords from voice information and detecting sudden changes in movement from gyro sensor information;
[1331] If an abnormality is detected based on the analysis results, an alert will be sent and, if necessary, a security alarm will be activated.
[1332] A system including:
[1333] (Claim 2)
[1334] 10. The system of claim 1, further comprising means for analyzing the audio information and the gyro sensor information with a machine learning model.
[1335] (Claim 3)
[1336] 2. The system according to claim 1, further comprising means for determining, when an abnormality is detected, a destination of an alert depending on the urgency of the abnormality.
[1337] "Example 2: Combining Emotion Engines"
[1338] (Claim 1)
[1339] A means for acquiring voice information;
[1340] A means for acquiring gyro sensor information;
[1341] means for transmitting the acquired voice information and gyro sensor information to a server;
[1342] A means for the server to analyze the audio information and the gyro sensor information to detect an abnormality;
[1343] A means for sending an alert to the user and their relatives when an anomaly is detected;
[1344] A means of sending a notification to the police when a serious abnormality is detected;
[1345] A means for activating a security alarm when an abnormality is detected;
[1346] Includes an emotion engine that recognizes user emotions,
[1347] A means for transmitting the acquired voice information to an emotion engine to analyze the user's emotion;
[1348] means for transmitting the analyzed emotion information to a server;
[1349] A system including:
[1350] (Claim 2)
[1351] 10. The system of claim 1, further comprising means for analyzing emotion information together with audio information and gyro sensor information using a machine learning model.
[1352] (Claim 3)
[1353] 2. The system according to claim 1, further comprising means for determining, when an abnormality is detected, a destination of an alert based on the urgency of the abnormality and the user's emotion.
[1354] "Application example 2 when combining emotion engines"
[1355] (Claim 1)
[1356] A means for acquiring voice information;
[1357] A means for acquiring gyro sensor information;
[1358] means for transmitting the acquired voice information and gyro sensor information to a server;
[1359] A means for the server to analyze the audio information and the gyro sensor information to detect an abnormality;
[1360] A means for sending an alert to the user and their relatives when an anomaly is detected;
[1361] A means of sending a notification to the police when a serious abnormality is detected;
[1362] A means for activating a security alarm when an abnormality is detected;
[1363] means including an emotion engine for recognizing an emotion of a user;
[1364] means for transmitting the emotion information recognized by the emotion engine to a server;
[1365] A means for supporting the safety of users using smart glasses;
[1366] A system including:
[1367] (Claim 2)
[1368] 10. The system of claim 1, further comprising means for analyzing the audio information, the gyro sensor information, and the emotion information using a machine learning model.
[1369] (Claim 3)
[1370] 2. The system according to claim 1, further comprising means for determining an alert destination in accordance with the urgency of an abnormality when the abnormality is detected, and controlling the activation of an appropriate alert and a security buzzer. [Explanation of symbols]
[1371] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring voice information; A means for acquiring gyro sensor information; means for transmitting the acquired voice information and gyro sensor information to a server; A means for the server to analyze the audio information and the gyro sensor information to detect an abnormality; A means for sending an alert to the user and their relatives when an anomaly is detected; A means of sending a notification to the police when a serious abnormality is detected; A means for activating a security alarm when an abnormality is detected; A system including:
2. The system of claim 1 , further comprising means for analyzing the audio information and the gyro sensor information with a machine learning model.
3. 2. The system according to claim 1, further comprising means for determining, when an abnormality is detected, a destination of an alert depending on the urgency of the abnormality.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A