system
The system addresses the challenge of rapid emergency detection and response in elderly care settings by using voice and environmental sensors, AI analysis, and communication, ensuring timely assistance and adaptive system enhancements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional systems lack the capability to quickly detect and respond to emergencies in the living environment of elderly individuals, particularly in situations where immediate assistance is required, such as solitary deaths, due to insufficient means for voice detection and environmental anomaly recognition.
A system incorporating a voice recognition sensor for detecting specific keywords, environmental data sensors for anomaly detection, an artificial intelligence module for analysis, and communication means for immediate notifications, with user feedback mechanisms to enhance system accuracy.
Enables rapid emergency response by identifying critical situations and sending notifications to relevant authorities, while allowing for continuous system improvement through user feedback.
Smart Images

Figure 2026070287000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present invention is to provide a system that can quickly detect and respond to unexpected situations and the risk of solitary death in the living environment of the elderly, single elderly people, etc. Conventional technologies have a problem that means for immediately responding in an emergency are insufficient, and as a result, it is difficult to ensure safety. For example, means for detecting a voice asking for help and promptly responding thereto are required.
Means for Solving the Problems
[0005] This invention solves these problems with a system that includes a voice recognition sensor that detects predetermined keywords and a sensor that collects environmental data for detecting anomalies. The system includes a communication means that analyzes the collected data with an artificial intelligence module and generates and transmits notifications based on the analysis results. This configuration allows for immediate emergency notifications and strengthens safety measures. Furthermore, by receiving feedback from users and adjusting system parameters, even more effective measures can be implemented.
[0006] A "voice recognition sensor" is an electronic device used to detect specific sounds or keywords.
[0007] "Environmental data" refers to information collected by sensors, such as ambient temperature, light intensity, and motion.
[0008] An "artificial intelligence module" is a program or system that analyzes collected data and detects anomalies.
[0009] "Communication means" refers to a function or device for sending generated notifications or messages to other devices or systems.
[0010] "System parameters" are variables or values that control the operating characteristics of a system and allow for adjustment of its settings.
[0011] "Feedback" refers to opinions and information provided by users with the aim of improving the system. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] In an embodiment for carrying out the present invention, the system is configured as follows: This system includes a voice recognition sensor, a sensor for collecting environmental data, an artificial intelligence module, and communication means, all of which work together.
[0034] First, the device uses a voice recognition sensor to detect specific keywords, such as "Help!!". This automatically starts data collection in emergencies. In addition, sensors that collect environmental data continuously monitor surrounding environmental information such as temperature, movement, and light intensity. This provides a foundation for detecting abnormal situations that would not normally occur.
[0035] Next, the server receives the data collected from these sensors and analyzes it using an artificial intelligence module. This analysis involves pattern recognition for anomaly detection and comparison with historical data to determine any abnormalities. If an anomaly is detected, the server generates an emergency notification and transmits it to property managers and relevant authorities via communication channels.
[0036] Users receive notifications and take the necessary actions according to the instructions provided. In some cases, they can also contribute to system improvements by sending feedback to the server. Based on this feedback, the server adjusts system parameters to improve the accuracy of anomaly detection in the future.
[0037] For example, if a terminal detects the cry "Help!!", the audio data is immediately sent to a server for analysis. If it is determined that there is an emergency, such as someone having collapsed and being unable to move, the administrator is immediately notified. Through this series of processes, the system is structured to enable a rapid response to protect the safety of residents.
[0038] In this way, the present invention provides an embodiment that allows for a comprehensive understanding of the surrounding situation and immediate response in the event of an emergency.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The device activates a voice recognition sensor to constantly monitor surrounding sounds and detects the keyword "Help!!". The voice recognition sensor analyzes specific voice patterns in real time and collects information when the keyword is detected.
[0042] Step 2:
[0043] The device activates sensors to collect environmental data in response to detection by the voice recognition sensor, recording information about the surrounding environment such as temperature, light intensity, and motion. This data is an important element for anomaly detection.
[0044] Step 3:
[0045] The device transmits collected audio and environmental data to the server. The data includes audio clips, various sensor information, and time information.
[0046] Step 4:
[0047] The server receives data sent from the terminal and stores it in the database. It then starts analyzing the received data using an artificial intelligence module.
[0048] Step 5:
[0049] The server uses an artificial intelligence module to analyze data, checking for the presence or absence of keywords in audio clips and for abnormal patterns in environmental data. If an anomaly is detected as a result of the analysis, an action is triggered.
[0050] Step 6:
[0051] When an anomaly is detected, the server determines what information to send in the notification and generates a notification message. This message includes the type of anomaly, the time it occurred, and the recommended action.
[0052] Step 7:
[0053] The server sends the generated notification message to the terminal of a pre-registered property manager or related organization. Notifications are sent via email or SMS using various communication methods.
[0054] Step 8:
[0055] Users receive notifications from the server on their devices and take necessary actions based on the content. For example, they might take emergency action to verify the site or ensure safety.
[0056] Step 9:
[0057] Users send feedback to the server regarding the results of their actions and areas for improvement. The server uses the received feedback to adjust system parameters and improve analysis accuracy.
[0058] (Example 1)
[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0060] In recent years, there has been a growing demand for improved security and safety. However, conventional systems have a problem in that functions such as speech recognition, environmental data collection, and anomaly detection operate individually, making it difficult to respond quickly to abnormal situations. Furthermore, they lack mechanisms to improve system accuracy by utilizing user feedback. Therefore, there is a need to realize a system that can respond effectively and quickly to emergencies with a more integrated approach.
[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0062] In this invention, the server includes voice recognition means for detecting predetermined keywords, means for collecting environmental information to monitor the surrounding conditions and detect anomalies, analysis means equipped with an intelligent module for processing the generated data, communication means for creating and transmitting warnings based on the analysis results, and means for receiving information from the user and adjusting system settings. This not only enables a rapid and accurate response to emergencies, but also allows for continuous improvement of the system's accuracy through user feedback.
[0063] A "designated keyword" is a specific word or phrase that has been set in advance and acts as a linguistic trigger that, when detected by speech recognition, causes a specific action.
[0064] "Voice recognition means" refers to a technology or device for sensing surrounding sounds and identifying specific keywords or phrases.
[0065] "Means of collecting environmental information" refers to devices and technologies that have the function of continuously monitoring the surrounding physical conditions, such as temperature, movement, and light intensity, and collecting them as data.
[0066] An "analysis means equipped with an intelligent module" is a computer system that includes algorithms or programs for performing learning and inference based on collected data, and for performing anomaly detection and pattern recognition.
[0067] "Communication methods" refer to technologies and devices that can send and receive data and warnings to remote locations, and typically utilize the internet or wireless communication.
[0068] "Means of receiving information from users and adjusting system settings" refers to technologies that allow a system to accept user feedback and requests and automatically modify or optimize its operation and parameters based on that information.
[0069] This invention provides an integrated system for detecting specific keywords, immediately recognizing abnormal environmental conditions, and taking appropriate action. Its configuration and operation are described below.
[0070] The device is equipped with a voice recognition sensor that constantly monitors the surrounding audio environment. This voice recognition sensor uses a high-performance microphone and voice analysis software to quickly identify situations requiring attention by detecting predetermined keywords, such as "Help!!".
[0071] Furthermore, the device has multiple environmental information gathering devices, including temperature, motion, and light sensors. These sensors continuously record the surrounding environmental conditions, providing a foundation for detecting signs of abnormal events or situations.
[0072] The server aggregates data sent from terminals and performs detailed analysis using an artificial intelligence module. This AI module has the ability to recognize anomalous patterns by comparing them with past data using machine learning algorithms. Based on the analysis results, the server can determine if an anomaly has occurred and immediately create an emergency notification.
[0073] Notifications generated through communication channels are immediately sent to property managers and relevant authorities. This enables a prompt response and protects the safety of residents.
[0074] Users take swift action based on received notifications. Depending on the situation, users can send feedback to the server. This feedback is used to optimize the system and improve the accuracy of anomaly detection in the future.
[0075] For example, if someone shouts "Help!!" in the hallway of an apartment building, the audio data is immediately sent to a server. If the audio analysis confirms that it is an urgent situation, a notification is sent to the administrator urging them to take immediate action.
[0076] A concrete example of a prompt for a generating AI model would be, "Please provide detailed information on the functions of sensors and communication systems for detecting and responding quickly to potential emergencies that may occur in real estate facilities." Based on this prompt, the generating AI would propose effective countermeasures appropriate to the situation.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The device constantly monitors ambient sounds using a voice recognition sensor. It receives ambient sound data as input and performs data processing to analyze whether it contains specific keywords. Specifically, the device performs frequency and phoneme analysis on the audio to detect the keyword "Help!!". Once the keyword is detected, it generates an alert and proceeds to the next step.
[0080] Step 2:
[0081] The device collects input data from temperature sensors, motion sensors, and light sensors. Based on this data, data processing is performed to identify unusual patterns. Specifically, real-time data from the sensors is integrated, and an algorithm is applied to detect anomalies. When an abnormal environmental change is detected, the aggregated data is sent to the server.
[0082] Step 3:
[0083] The server receives voice data and environmental data transmitted from the terminal. The input includes voice and environmental sensor data, and data calculations are performed using an artificial intelligence module to analyze this data. Specifically, anomaly patterns are recognized using a machine learning model and compared with past data to calculate the probability of an anomaly. If an anomaly is detected as a result of the analysis, the next step is to generate an emergency notification.
[0084] Step 4:
[0085] The server initiates a process to create notification content based on the analysis results. The input is the analyzed anomaly identification information, which is used to construct the notification message. Specifically, it automatically generates an emergency notification containing details of the situation for administrators or relevant authorities. Once the notification content is finalized, it is sent to the recipient via communication means.
[0086] Step 5:
[0087] The user receives a notification from the server. Based on the notification content, the user understands the situation and prepares to follow specific action instructions. Specifically, the user who checks the notification is expected to verify the situation on-site, take necessary actions, and send feedback to the server as needed, thereby contributing to improvements in the future.
[0088] (Application Example 1)
[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0090] In recent years, public and commercial facilities have been required to respond quickly to emergencies. However, these facilities need to cover a wide area, making it difficult to continuously monitor all emergencies using only human resources. Therefore, there is a need for a system that can instantly analyze various environmental information, detect anomalies, and quickly notify relevant parties.
[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0092] In this invention, the server includes means for a speech recognition detection device to sense specific instructional phrases, a sensor device for measuring numerical values of the external environment, and a machine learning module means for evaluating the acquired information. This makes it possible to quickly and accurately detect emergencies within public facilities and immediately notify relevant parties.
[0093] A "speech recognition detection device" is a device that detects specific instructional words as sounds and converts them into digital data.
[0094] A "sensor device" is a device that measures the physical characteristics and conditions of the external environment (e.g., temperature, light, motion) and collects them as data.
[0095] A "machine learning module" is a system consisting of algorithms and their implementations that use acquired information to detect patterns and anomalies and perform analysis.
[0096] A "communication device" is a device used to transmit notifications generated based on the analysis results to relevant terminals and equipment.
[0097] A "notification function" is a function that, when an emergency or abnormality is detected, sends appropriate notifications to pre-configured relevant parties or systems.
[0098] In an embodiment of this invention, the system is configured as follows: A server uses a speech recognition detection device to sense specific instructional phrases, which triggers a sensor device to measure numerical values of the external environment. The measured data is evaluated by a machine learning module to analyze whether or not there are any anomalies.
[0099] Specifically, the system uses the Google® Speech Recognition API as a speech recognition detection device, converting specific instructional phrases into digital data when they are spoken. This resulting speech data, along with environmental data simultaneously acquired from sensor devices, is sent to a machine learning module for pattern recognition and anomaly detection.
[0100] Furthermore, the server generates a notification based on the analysis results via a communication device and quickly sends the notification to registered terminals. The notification function is activated, informing relevant parties of the emergency.
[0101] For example, if a voice shouting "Help!!" is detected inside a shopping mall, the footage and temperature data from the scene are analyzed, and a notification is sent to security personnel as needed. This notification includes detailed instructions for action, enabling a swift emergency response.
[0102] An example of a prompt to input into the generating AI model would be: "A voice suddenly shouting 'Help!!' was detected inside a shopping mall. The on-site video and temperature data are as follows... Based on this, please provide detailed instructions on how employees should respond."
[0103] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0104] Step 1:
[0105] Start voice recognition
[0106] The device initiates speech recognition using the Google Speech Recognition API. The input is ambient sound, and the device detects whether it contains the specific demonstrative phrase "Help!!". If detected, the audio data is sent to the server as digital data.
[0107] Step 2:
[0108] Environmental data collection
[0109] The server measures external environmental data from sensor devices in parallel with speech recognition. These measurements include temperature sensors and motion detection sensors, and the collected data includes temperature, motion, and light intensity. The input here is the sensor values, and the output is the aggregated environmental data.
[0110] Step 3:
[0111] Data Analysis
[0112] The server sends the acquired audio and environmental data to a machine learning module. Based on the input data, the module analyzes the data to determine if it is an anomaly by comparing it to past patterns. The output is the result of the anomaly determination. Specifically, if there is a possibility of an anomaly, the module also calculates the confidence level of that determination.
[0113] Step 4:
[0114] Notification generation
[0115] If an anomaly is detected, the server generates a notification via the communication device. The input is the result of the anomaly analysis, and the output is an emergency notification message. Specifically, the message includes a description of the situation on site and recommended actions.
[0116] Step 5:
[0117] Send notification
[0118] The server sends generated notifications to registered devices. The input is the notification message, and the output is the notification displayed on the device. This allows users to confirm the details of an emergency and take prompt action. The notification provides specific action guidelines, ensuring that appropriate instructions are conveyed to employees and relevant parties.
[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0120] In an embodiment for carrying out the present invention, the system comprises a voice recognition sensor, an environmental data collection sensor, an artificial intelligence module, a communication means, and an emotion recognition engine.
[0121] First, the device uses a voice recognition sensor to continuously monitor for emergency keywords such as "Help!!". In addition, an emotion recognition engine analyzes the user's voice and surrounding sounds to determine their emotional state. This analysis makes it possible to understand whether the user is stressed or calm.
[0122] Next, the device sends the collected data to the server. The server processes the received data using an artificial intelligence module. This process combines voice and environmental data for detailed analysis to understand the user's situation. Emotional state data provided by the emotion recognition engine is also used as part of the analysis. This allows for the detection of abnormal conditions that require immediate attention.
[0123] When an anomaly is detected, the server generates a notification depending on the situation. At this time, emotional information obtained from the emotion recognition engine is used to customize the notification message, resulting in a more appropriate notification. For example, if the user is experiencing extreme anxiety or fear, this will be reflected in the notification, prompting the administrator receiving the notification to take prompt action.
[0124] Users receive notifications from the server and take action based on the situation on-site. The notification recipients can immediately understand the urgency and severity of the situation through messages that include emotional information. Feedback received after the response is sent back to the server, which is used to help the system learn and improve.
[0125] For example, if the emergency keyword "Help!!" is recognized, and the emotion recognition engine determines from the tone of voice that the user is in a state of extreme fear, a notification is generated more quickly than usual, prompting immediate action. Based on this information, the server sends an emergency response alert to the property manager.
[0126] In this way, the system, which incorporates an emotion recognition engine, has a structure that enables more appropriate and rapid responses by taking the user's emotional state into consideration during decision-making.
[0127] The following describes the processing flow.
[0128] Step 1:
[0129] The device constantly monitors ambient sounds using a voice recognition sensor to detect whether "Help!!" or other emergency keywords are spoken. Additionally, an emotion engine analyzes the user's voice tone to recognize emotional states such as stress or anxiety.
[0130] Step 2:
[0131] The device transmits identified voice data and emotional state data to the server. Simultaneously, data such as temperature, light intensity, and motion collected by environmental sensors are also transferred to the server. This provides comprehensive information about the user's state and living environment.
[0132] Step 3:
[0133] The server receives data sent from the terminal and stores it in a database. This data is analyzed using an artificial intelligence module to check for abnormalities in voice keywords and emotional states.
[0134] Step 4:
[0135] The server determines whether an anomaly has been detected based on the analysis results, including the output of the emotion engine. If the emotional state is particularly unstable, it is more likely to be treated as an anomaly.
[0136] Step 5:
[0137] If an anomaly is detected, the server generates a notification message tailored to the nature of the anomaly. This message is customized based on the user's emotional state, with its wording adjusted based on the level of urgency and the corresponding emotion.
[0138] Step 6:
[0139] The server sends the generated notification message to the terminal of the registered property manager or relevant organization. This is done using standard communication methods such as email or SMS.
[0140] Step 7:
[0141] Users receive notifications from the server and take prompt action based on their content. These notifications include the cause of the anomaly and the user's estimated emotional state, allowing for a more accurate understanding of the situation.
[0142] Step 8:
[0143] After completing the task, users send feedback to the server via their device regarding the results and any discoveries they made. This feedback is used as training data for the system and contributes to future system improvements.
[0144] (Example 2)
[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0146] Conventional speech recognition systems primarily rely on keyword detection for anomaly detection, making it difficult to provide rapid and appropriate responses that take into account the user's emotional state. Furthermore, they lacked mechanisms to comprehensively analyze environmental information to determine urgency and provide appropriate notifications. Additionally, they lacked effective means of utilizing user feedback for system improvement.
[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0148] In this invention, the server includes means for a speech recognition device to monitor specified keywords, a collection device for collecting environmental information and detecting anomalies, a machine learning module for evaluating the analyzed data in detail, and an emotion recognition engine for determining and utilizing emotional states in analysis. This makes it possible to comprehensively analyze the user's emotional state and environmental information and generate quick and appropriate notifications. Furthermore, user feedback can be used to optimize the system.
[0149] A "voice recognition device" is a device that continuously monitors surrounding sounds and has the function of detecting specific keywords.
[0150] A "collection device" is a device used to collect information about the surrounding environment and detect anomalies.
[0151] A "machine learning module" is a system that includes algorithms and models for performing detailed analysis on collected data and evaluating the information.
[0152] A "communication device" is a device that generates notifications based on analysis results and transmits them externally.
[0153] An "emotion recognition engine" refers to a technology that identifies emotional states from voice data and uses that information for analysis.
[0154] This system includes a speech recognition device, data collection device, machine learning module, emotion recognition engine, and communication device for the user. The terminal uses the speech recognition device to acquire voice data from the user's surroundings. This device has the function of continuously monitoring specific keywords and detecting keywords for emergencies.
[0155] Furthermore, the terminal is equipped with a data collection device for collecting environmental data. This device is intended to collect information about the user's surroundings and detect anomalies. The collected data is transmitted from the terminal to a server.
[0156] The server uses a machine learning module to analyze voice data and environmental information in detail. This module comprehensively evaluates the data based on a generative AI model to determine the user's situation. It also uses an emotion recognition engine to identify the user's emotional state from their voice and incorporates this into the analysis results.
[0157] For example, if the keyword "Help!!" is detected and strong fear is perceived from the user's voice, this information will be included in the analysis results. Based on the analysis results, the server will use a communication device to automatically generate an appropriate notification and send it to the administrator.
[0158] Upon receiving a notification, the user takes action based on the content of the notification, according to the situation at hand. An example of a prompt message, used when analyzing a specific emotional state, is, "Analyze the emotional state of this audio sample and determine whether it is an emergency." This allows the system to ensure the user's safety in real time and enable a rapid response.
[0159] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0160] Step 1:
[0161] The device acquires surrounding audio in real time using a speech recognition device. The input is raw audio data from the surroundings. The speech recognition device processes this data and detects a specified keyword, such as "Help!!". The output is a flag indicating that the keyword was detected. Specifically, when audio is input, a high-speed speech analysis algorithm is applied to determine the presence or absence of the keyword.
[0162] Step 2:
[0163] The device uses an emotion recognition engine to perform emotion analysis on the audio data acquired in Step 1. The input is audio data, and the emotion recognition engine analyzes the tone, speed, and volume of the voice to determine the user's emotional state. The output is data indicating the emotional state. For example, if the voice contains clear signs of anxiety or fear, that emotion will be detected.
[0164] Step 3:
[0165] The terminal uses a data collection device to gather information about the surrounding environment. The input is environmental data from sensors. The data collection device acquires information such as temperature, humidity, and brightness, and checks for any abnormal values. The output is a set of current environmental information. For example, if there is an abnormal temperature rise, that information is recorded.
[0166] Step 4:
[0167] The terminal sends the data obtained in steps 1, 2, and 3 to the server. The input consists of voice keyword flags, emotional state data, and environmental information. This data is sent to the server via the internet. The output is a signal confirming that this data has arrived at the server. A secure protocol is used to ensure the security of the data.
[0168] Step 5:
[0169] The server uses a machine learning module to analyze data sent from the terminal in detail. Inputs include voice keyword flags, emotional state data, and environmental information. The machine learning module uses a generative AI model to comprehensively evaluate the data and determine the user's situation. The output is an evaluation result indicating the user's situation. For example, if an emergency is determined, that information will be generated.
[0170] Step 6:
[0171] The server generates notifications based on the analysis results and sends the appropriate notifications to the administrator using a communication device. The input is the server's evaluation result. A dedicated template is used to generate the content of the notification. The output is the generated notification message. For example, a notification indicating an emergency is sent to the administrator.
[0172] Step 7:
[0173] The user receives notifications from the server and takes appropriate action on-site based on the notification content. The input is the notification message sent from the server. The user understands its content and takes specific action. The output is the user's action. If feedback information is obtained, it is also sent to the server.
[0174] (Application Example 2)
[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0176] Current security systems lack sufficient emergency response capabilities that take into account the user's emotional state, making it difficult to immediately assess the severity and urgency of an abnormal situation. Therefore, appropriate and rapid responses tailored to the situation are required.
[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0178] In this invention, the server includes means for a speech recognition device to detect predetermined emergency terms, means for analyzing the user's emotional state using an emotion recognition engine, and a device for collecting environmental information. This makes it possible to determine the urgency of an abnormal situation based on the user's emotional state and respond immediately.
[0179] A "speech recognition device" is a device that converts speech into digital signals and has the function of detecting specific keywords or phrases in real time.
[0180] An "emotion recognition engine" is an engine equipped with an algorithm that analyzes a user's emotional state from voice data and environmental data and identifies specific emotions.
[0181] An "environmental information collecting device" is a device that acquires ambient environmental data such as temperature, light, and sound, and provides that data as basic information for analysis.
[0182] The "intelligent module function" is an artificial intelligence analysis function that comprehensively evaluates speech recognition, emotion recognition, and environmental information to determine the presence and urgency of abnormal situations.
[0183] The "communication function" refers to the function that performs digital communication to send notifications generated based on the analysis results to a designated terminal.
[0184] The system implementing this invention is comprised of a combination of a speech recognition device, an emotion recognition engine, a device for collecting environmental information, an intelligent module function, and a communication function. The server uses the speech recognition device to analyze voice commands from the user and detect specific emergency terms in real time. In this process, the voice data is captured as a digital signal using the SpeechRecognition library. Furthermore, the emotion recognition engine using TENSORFLOW® analyzes the user's emotional state from the tone of the voice and ambient acoustics.
[0185] Meanwhile, the terminal acquires data such as temperature, light, and sound through a device that collects environmental information. The intelligent module function on the server comprehensively analyzes this data to determine whether there is an abnormal situation and how urgent it is. For example, if an unusually loud sound is detected at night and the user's emotions indicate strong anxiety or fear, the system will use that information to determine that a rapid response is necessary.
[0186] As a result, communication functions are used to send emergency notifications to relevant devices based on the results of analysis and sentiment evaluation. This allows appropriate administrators and stakeholders to quickly understand the situation and take necessary actions.
[0187] For example, if a user suddenly shouts "Help!!" at home one night, the system will instantly recognize the voice, and its emotion recognition engine will then identify any signs of anxiety. The server will analyze the situation and, if necessary, send a notification to the security administrator stating, "Unusual noise and strong anxiety in the middle of the night; immediate action required."
[0188] Examples of prompt messages include, "How does the system respond when an abnormal sound is detected?" and "What is the procedure when strong feelings of anxiety are detected?"
[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0190] Step 1:
[0191] The terminal uses its built-in microphone to capture the user's voice and inputs it into a speech recognition device. The speech recognition device uses the SpeechRecognition library to convert the voice into text data and obtains an output that determines whether or not it contains emergency terms.
[0192] Step 2:
[0193] The device sends voice data and ambient sound environment data to the emotion recognition engine. The emotion recognition engine uses TensorFlow to analyze the tone and patterns of the voice and identify the user's emotional state. This process outputs an emotion classification result, such as fear or anxiety.
[0194] Step 3:
[0195] The server collects data acquired from devices that gather environmental information. Digital data such as temperature, light, and ambient noise levels are input into an analysis program to evaluate whether the environment is in an abnormal state. If an abnormal value is detected, the evaluation result is output.
[0196] Step 4:
[0197] The intelligent module function on the server comprehensively analyzes the speech recognition results and emotional state obtained from steps 1 and 2, and the environmental evaluation data from step 3. By correlating each input data, an analysis result is obtained that determines whether an abnormal situation exists and its level of urgency.
[0198] Step 5:
[0199] The server generates notifications via communication functions based on the analysis results of abnormal situations and urgency levels from the intelligent module. These notifications include the user's emotional state and environmental information, and are sent to the relevant administrator terminal as emergency notifications when a rapid response is required.
[0200] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0201] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0202] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0203] [Second Embodiment]
[0204] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0205] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0206] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0207] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0208] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0209] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0210] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0211] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0212] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0213] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0214] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0215] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0216] In an embodiment for carrying out the present invention, the system is configured as follows: This system includes a voice recognition sensor, a sensor for collecting environmental data, an artificial intelligence module, and communication means, all of which work together.
[0217] First, the device uses a voice recognition sensor to detect specific keywords, such as "Help!!". This automatically starts data collection in emergencies. In addition, sensors that collect environmental data continuously monitor surrounding environmental information such as temperature, movement, and light intensity. This provides a foundation for detecting abnormal situations that would not normally occur.
[0218] Next, the server receives the data collected from these sensors and analyzes it using an artificial intelligence module. This analysis involves pattern recognition for anomaly detection and comparison with historical data to determine any abnormalities. If an anomaly is detected, the server generates an emergency notification and transmits it to property managers and relevant authorities via communication channels.
[0219] Users receive notifications and take the necessary actions according to the instructions provided. In some cases, they can also contribute to system improvements by sending feedback to the server. Based on this feedback, the server adjusts system parameters to improve the accuracy of anomaly detection in the future.
[0220] For example, if a terminal detects the cry "Help!!", the audio data is immediately sent to a server for analysis. If it is determined that there is an emergency, such as someone having collapsed and being unable to move, the administrator is immediately notified. Through this series of processes, the system is structured to enable a rapid response to protect the safety of residents.
[0221] In this way, the present invention provides an embodiment that allows for a comprehensive understanding of the surrounding situation and immediate response in the event of an emergency.
[0222] The following describes the processing flow.
[0223] Step 1:
[0224] The device activates a voice recognition sensor to constantly monitor surrounding sounds and detects the keyword "Help!!". The voice recognition sensor analyzes specific voice patterns in real time and collects information when the keyword is detected.
[0225] Step 2:
[0226] The device activates sensors to collect environmental data in response to detection by the voice recognition sensor, recording information about the surrounding environment such as temperature, light intensity, and motion. This data is an important element for anomaly detection.
[0227] Step 3:
[0228] The device transmits collected audio and environmental data to the server. The data includes audio clips, various sensor information, and time information.
[0229] Step 4:
[0230] The server receives data sent from the terminal and stores it in the database. It then starts analyzing the received data using an artificial intelligence module.
[0231] Step 5:
[0232] The server uses an artificial intelligence module to analyze data, checking for the presence or absence of keywords in audio clips and for abnormal patterns in environmental data. If an anomaly is detected as a result of the analysis, an action is triggered.
[0233] Step 6:
[0234] When an anomaly is detected, the server determines what information to send in the notification and generates a notification message. This message includes the type of anomaly, the time it occurred, and the recommended action.
[0235] Step 7:
[0236] The server sends the generated notification message to the terminal of a pre-registered property manager or related organization. Notifications are sent via email or SMS using various communication methods.
[0237] Step 8:
[0238] Users receive notifications from the server on their devices and take necessary actions based on the content. For example, they might take emergency action to verify the site or ensure safety.
[0239] Step 9:
[0240] Users send feedback to the server regarding the results of their actions and areas for improvement. The server uses the received feedback to adjust system parameters and improve analysis accuracy.
[0241] (Example 1)
[0242] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0243] In recent years, there has been a growing demand for improved security and safety. However, conventional systems have a problem in that functions such as speech recognition, environmental data collection, and anomaly detection operate individually, making it difficult to respond quickly to abnormal situations. Furthermore, they lack mechanisms to improve system accuracy by utilizing user feedback. Therefore, there is a need to realize a system that can respond effectively and quickly to emergencies with a more integrated approach.
[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0245] In this invention, the server includes voice recognition means for detecting predetermined keywords, means for collecting environmental information to monitor the surrounding conditions and detect anomalies, analysis means equipped with an intelligent module for processing the generated data, communication means for creating and transmitting warnings based on the analysis results, and means for receiving information from the user and adjusting system settings. This not only enables a rapid and accurate response to emergencies, but also allows for continuous improvement of the system's accuracy through user feedback.
[0246] A "designated keyword" is a specific word or phrase that has been set in advance and acts as a linguistic trigger that, when detected by speech recognition, causes a specific action.
[0247] "Voice recognition means" refers to a technology or device for sensing surrounding sounds and identifying specific keywords or phrases.
[0248] "Means of collecting environmental information" refers to devices and technologies that have the function of continuously monitoring the surrounding physical conditions, such as temperature, movement, and light intensity, and collecting them as data.
[0249] An "analysis means equipped with an intelligent module" is a computer system that includes algorithms or programs for performing learning and inference based on collected data, and for performing anomaly detection and pattern recognition.
[0250] "Communication methods" refer to technologies and devices that can send and receive data and warnings to remote locations, and typically utilize the internet or wireless communication.
[0251] "Means of receiving information from users and adjusting system settings" refers to technologies that allow a system to accept user feedback and requests and automatically modify or optimize its operation and parameters based on that information.
[0252] This invention provides an integrated system for detecting specific keywords, immediately recognizing abnormal environmental conditions, and taking appropriate action. Its configuration and operation are described below.
[0253] The device is equipped with a voice recognition sensor that constantly monitors the surrounding audio environment. This voice recognition sensor uses a high-performance microphone and voice analysis software to quickly identify situations requiring attention by detecting predetermined keywords, such as "Help!!".
[0254] Furthermore, the device has multiple environmental information gathering devices, including temperature, motion, and light sensors. These sensors continuously record the surrounding environmental conditions, providing a foundation for detecting signs of abnormal events or situations.
[0255] The server aggregates data sent from terminals and performs detailed analysis using an artificial intelligence module. This AI module has the ability to recognize anomalous patterns by comparing them with past data using machine learning algorithms. Based on the analysis results, the server can determine if an anomaly has occurred and immediately create an emergency notification.
[0256] Notifications generated through communication channels are immediately sent to property managers and relevant authorities. This enables a prompt response and protects the safety of residents.
[0257] Users take swift action based on received notifications. Depending on the situation, users can send feedback to the server. This feedback is used to optimize the system and improve the accuracy of anomaly detection in the future.
[0258] For example, if someone shouts "Help!!" in the hallway of an apartment building, the audio data is immediately sent to a server. If the audio analysis confirms that it is an urgent situation, a notification is sent to the administrator urging them to take immediate action.
[0259] A concrete example of a prompt for a generating AI model would be, "Please provide detailed information on the functions of sensors and communication systems for detecting and responding quickly to potential emergencies that may occur in real estate facilities." Based on this prompt, the generating AI would propose effective countermeasures appropriate to the situation.
[0260] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0261] Step 1:
[0262] The device constantly monitors ambient sounds using a voice recognition sensor. It receives ambient sound data as input and performs data processing to analyze whether it contains specific keywords. Specifically, the device performs frequency and phoneme analysis on the audio to detect the keyword "Help!!". Once the keyword is detected, it generates an alert and proceeds to the next step.
[0263] Step 2:
[0264] The device collects input data from temperature sensors, motion sensors, and light sensors. Based on this data, data processing is performed to identify unusual patterns. Specifically, real-time data from the sensors is integrated, and an algorithm is applied to detect anomalies. When an abnormal environmental change is detected, the aggregated data is sent to the server.
[0265] Step 3:
[0266] The server receives voice data and environmental data transmitted from the terminal. The input includes voice and environmental sensor data, and data calculations are performed using an artificial intelligence module to analyze this data. Specifically, anomaly patterns are recognized using a machine learning model and compared with past data to calculate the probability of an anomaly. If an anomaly is detected as a result of the analysis, the next step is to generate an emergency notification.
[0267] Step 4:
[0268] The server initiates a process to create notification content based on the analysis results. The input is the analyzed anomaly identification information, which is used to construct the notification message. Specifically, it automatically generates an emergency notification containing details of the situation for administrators or relevant authorities. Once the notification content is finalized, it is sent to the recipient via communication means.
[0269] Step 5:
[0270] The user receives a notification from the server. Based on the notification content, the user understands the situation and prepares to follow specific action instructions. Specifically, the user who checks the notification is expected to verify the situation on-site, take necessary actions, and send feedback to the server as needed, thereby contributing to improvements in the future.
[0271] (Application Example 1)
[0272] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0273] In recent years, public and commercial facilities have been required to respond quickly to emergencies. However, these facilities need to cover a wide area, making it difficult to continuously monitor all emergencies using only human resources. Therefore, there is a need for a system that can instantly analyze various environmental information, detect anomalies, and quickly notify relevant parties.
[0274] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0275] In this invention, the server includes means for a speech recognition detection device to sense specific instructional phrases, a sensor device for measuring numerical values of the external environment, and a machine learning module means for evaluating the acquired information. This makes it possible to quickly and accurately detect emergencies within public facilities and immediately notify relevant parties.
[0276] A "speech recognition detection device" is a device that detects specific instructional words as sounds and converts them into digital data.
[0277] A "sensor device" is a device that measures the physical characteristics and conditions of the external environment (e.g., temperature, light, motion) and collects them as data.
[0278] The "machine learning module" is a system consisting of an algorithm for detecting patterns and anomalies using the acquired information and its implementation for analysis.
[0279] The "communication device" is a device for transmitting a notification generated based on the analyzed result to related terminals and devices.
[0280] The "reporting function means" is a function for appropriately reporting to pre-set related persons or systems when detecting an emergency or an anomaly.
[0281] As a form for implementing this invention, the system is configured as follows. The server uses a voice recognition detection device to sense a specific instruction sentence, which triggers the sensor device to measure numerical values of the external environment. The measured data is evaluated by the machine learning module to analyze the presence or absence of anomalies.
[0282] Specifically, as the voice recognition detection device, Google Speech Recognition API is used to convert it into digital data when a specific instruction sentence is uttered. The voice data thus obtained, together with the environmental data simultaneously acquired from the sensor device, is sent to the machine learning module to perform pattern recognition and anomaly detection.
[0283] Furthermore, the server creates a notification based on the analysis result via the communication device and quickly sends a notification to the registered terminals. The reporting function means works to inform the related persons of the emergency.
[0284] As a specific example, when a voice shouting "Help!!" is detected inside the facility of a shopping mall, analysis is performed together with the on-site video data and temperature data, and a notification is sent to the security personnel as necessary. Since the notification at this time includes detailed response instructions, a prompt emergency response is possible.
[0285] Examples of prompt sentences input into the generative AI model include ones such as "A voice shouting 'Help!!' was suddenly detected inside a certain shopping mall. The on-site video and temperature data are as follows... Based on this, please show in detail the procedures for how employees should respond."
[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0287] Step 1:
[0288] Start voice recognition
[0289] The terminal starts voice recognition using the Google Speech Recognition API. The input here is ambient sound, and it senses whether it contains a specific directive sentence "Help!!". If detected, the voice data is transmitted to the server as digital data.
[0290] Step 2:
[0291] Collect environmental data
[0292] The server measures the numerical values of the external environment from the sensor device in parallel with voice recognition. This measurement includes temperature sensors and motion detection sensors, and the collected data includes temperature, movement, light intensity, etc. The input here is sensor values, and the output is environmental data obtained by aggregating these.
[0293] Step 3:
[0294] Data analysis
[0295] The server transmits the obtained voice data and environmental data to the machine learning module. Based on the input data, the module compares it with past patterns and performs data analysis on whether it is abnormal. The output is the judgment result of whether there is an abnormality. As a specific operation, when there is a possibility of an abnormality, the confidence level is also calculated.
[0296] Step 4:
[0297] Notification generation
[0298] If an anomaly is detected, the server generates a notification via the communication device. The input is the result of the anomaly analysis, and the output is an emergency notification message. Specifically, the message includes a description of the situation on site and recommended actions.
[0299] Step 5:
[0300] Send notification
[0301] The server sends generated notifications to registered devices. The input is the notification message, and the output is the notification displayed on the device. This allows users to confirm the details of an emergency and take prompt action. The notification provides specific action guidelines, ensuring that appropriate instructions are conveyed to employees and relevant parties.
[0302] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0303] In an embodiment for carrying out the present invention, the system comprises a voice recognition sensor, an environmental data collection sensor, an artificial intelligence module, a communication means, and an emotion recognition engine.
[0304] First, the device uses a voice recognition sensor to continuously monitor for emergency keywords such as "Help!!". In addition, an emotion recognition engine analyzes the user's voice and surrounding sounds to determine their emotional state. This analysis makes it possible to understand whether the user is stressed or calm.
[0305] Next, the terminal transmits the data collected to the server. The server processes the received data using an artificial intelligence module. In this process, detailed analysis is performed by combining voice and environmental data to understand the user's situation. The emotion state data provided by the emotion recognition engine is also utilized as part of the analysis. As a result, an abnormal state that requires immediate response is detected.
[0306] When an abnormality is detected, the server generates a notification according to the situation. At this time, the emotion information obtained from the emotion recognition engine is used to customize the notification message, enabling a more appropriate expression for the notification. For example, when the user is in a state of extreme anxiety or fear, this fact is reflected in the notification content, making it possible to prompt the recipient, the administrator, to take prompt action.
[0307] The user receives the notification from the server and takes actions based on the on-site situation. The recipient of the notification can immediately understand the urgency and severity of the situation from the message with emotion information added. The feedback obtained after the response is sent back to the server again, which is useful for the learning and adjustment of the system.
[0308] As a specific example, when an emergency keyword "Help!!" is recognized and the emotion recognition engine determines that the user is in a very strong state of fear from the tone of voice, a notification is generated more quickly than usual, prompting an immediate response. The server sends an emergency response alert to the real estate administrator based on this information.
[0309] In this way, the system incorporating the emotion recognition engine has a structure that enables a more appropriate and rapid response by taking into account the user's emotional state in the judgment.
[0310] The following explains the processing flow.
[0311] Step 1:
[0312] The device constantly monitors ambient sounds using a voice recognition sensor to detect whether "Help!!" or other emergency keywords are spoken. Additionally, an emotion engine analyzes the user's voice tone to recognize emotional states such as stress or anxiety.
[0313] Step 2:
[0314] The device transmits identified voice data and emotional state data to the server. Simultaneously, data such as temperature, light intensity, and motion collected by environmental sensors are also transferred to the server. This provides comprehensive information about the user's state and living environment.
[0315] Step 3:
[0316] The server receives data sent from the terminal and stores it in a database. This data is analyzed using an artificial intelligence module to check for abnormalities in voice keywords and emotional states.
[0317] Step 4:
[0318] The server determines whether an anomaly has been detected based on the analysis results, including the output of the emotion engine. If the emotional state is particularly unstable, it is more likely to be treated as an anomaly.
[0319] Step 5:
[0320] If an anomaly is detected, the server generates a notification message tailored to the nature of the anomaly. This message is customized based on the user's emotional state, with its wording adjusted based on the level of urgency and the corresponding emotion.
[0321] Step 6:
[0322] The server sends the generated notification message to the terminal of the registered property manager or relevant organization. This is done using standard communication methods such as email or SMS.
[0323] Step 7:
[0324] Users receive notifications from the server and take prompt action based on their content. These notifications include the cause of the anomaly and the user's estimated emotional state, allowing for a more accurate understanding of the situation.
[0325] Step 8:
[0326] After completing the task, users send feedback to the server via their device regarding the results and any discoveries they made. This feedback is used as training data for the system and contributes to future system improvements.
[0327] (Example 2)
[0328] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0329] Conventional speech recognition systems primarily rely on keyword detection for anomaly detection, making it difficult to provide rapid and appropriate responses that take into account the user's emotional state. Furthermore, they lacked mechanisms to comprehensively analyze environmental information to determine urgency and provide appropriate notifications. Additionally, they lacked effective means of utilizing user feedback for system improvement.
[0330] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0331] In this invention, the server includes means for a speech recognition device to monitor specified keywords, a collection device for collecting environmental information and detecting anomalies, a machine learning module for evaluating the analyzed data in detail, and an emotion recognition engine for determining and utilizing emotional states in analysis. This makes it possible to comprehensively analyze the user's emotional state and environmental information and generate quick and appropriate notifications. Furthermore, user feedback can be used to optimize the system.
[0332] A "voice recognition device" is a device that continuously monitors surrounding sounds and has the function of detecting specific keywords.
[0333] A "collection device" is a device used to collect information about the surrounding environment and detect anomalies.
[0334] A "machine learning module" is a system that includes algorithms and models for performing detailed analysis on collected data and evaluating the information.
[0335] A "communication device" is a device that generates notifications based on analysis results and transmits them externally.
[0336] An "emotion recognition engine" refers to a technology that identifies emotional states from voice data and uses that information for analysis.
[0337] This system includes a speech recognition device, data collection device, machine learning module, emotion recognition engine, and communication device for the user. The terminal uses the speech recognition device to acquire voice data from the user's surroundings. This device has the function of continuously monitoring specific keywords and detecting keywords for emergencies.
[0338] Furthermore, the terminal is equipped with a data collection device for collecting environmental data. This device is intended to collect information about the user's surroundings and detect anomalies. The collected data is transmitted from the terminal to a server.
[0339] The server uses a machine learning module to analyze voice data and environmental information in detail. This module comprehensively evaluates the data based on a generative AI model to determine the user's situation. It also uses an emotion recognition engine to identify the user's emotional state from their voice and incorporates this into the analysis results.
[0340] For example, if the keyword "Help!!" is detected and strong fear is perceived from the user's voice, this information will be included in the analysis results. Based on the analysis results, the server will use a communication device to automatically generate an appropriate notification and send it to the administrator.
[0341] Upon receiving a notification, the user takes action based on the content of the notification, according to the situation at hand. An example of a prompt message, used when analyzing a specific emotional state, is, "Analyze the emotional state of this audio sample and determine whether it is an emergency." This allows the system to ensure the user's safety in real time and enable a rapid response.
[0342] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0343] Step 1:
[0344] The device acquires surrounding audio in real time using a speech recognition device. The input is raw audio data from the surroundings. The speech recognition device processes this data and detects a specified keyword, such as "Help!!". The output is a flag indicating that the keyword was detected. Specifically, when audio is input, a high-speed speech analysis algorithm is applied to determine the presence or absence of the keyword.
[0345] Step 2:
[0346] The device uses an emotion recognition engine to perform emotion analysis on the audio data acquired in Step 1. The input is audio data, and the emotion recognition engine analyzes the tone, speed, and volume of the voice to determine the user's emotional state. The output is data indicating the emotional state. For example, if the voice contains clear signs of anxiety or fear, that emotion will be detected.
[0347] Step 3:
[0348] The terminal uses a data collection device to gather information about the surrounding environment. The input is environmental data from sensors. The data collection device acquires information such as temperature, humidity, and brightness, and checks for any abnormal values. The output is a set of current environmental information. For example, if there is an abnormal temperature rise, that information is recorded.
[0349] Step 4:
[0350] The terminal sends the data obtained in steps 1, 2, and 3 to the server. The input consists of voice keyword flags, emotional state data, and environmental information. This data is sent to the server via the internet. The output is a signal confirming that this data has arrived at the server. A secure protocol is used to ensure the security of the data.
[0351] Step 5:
[0352] The server uses a machine learning module to analyze data sent from the terminal in detail. Inputs include voice keyword flags, emotional state data, and environmental information. The machine learning module uses a generative AI model to comprehensively evaluate the data and determine the user's situation. The output is an evaluation result indicating the user's situation. For example, if an emergency is determined, that information will be generated.
[0353] Step 6:
[0354] The server generates notifications based on the analysis results and sends the appropriate notifications to the administrator using a communication device. The input is the server's evaluation result. A dedicated template is used to generate the content of the notification. The output is the generated notification message. For example, a notification indicating an emergency is sent to the administrator.
[0355] Step 7:
[0356] The user receives notifications from the server and takes appropriate action on-site based on the notification content. The input is the notification message sent from the server. The user understands its content and takes specific action. The output is the user's action. If feedback information is obtained, it is also sent to the server.
[0357] (Application Example 2)
[0358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0359] Current security systems lack sufficient emergency response capabilities that take into account the user's emotional state, making it difficult to immediately assess the severity and urgency of an abnormal situation. Therefore, appropriate and rapid responses tailored to the situation are required.
[0360] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0361] In this invention, the server includes means for a speech recognition device to detect predetermined emergency terms, means for analyzing the user's emotional state using an emotion recognition engine, and a device for collecting environmental information. This makes it possible to determine the urgency of an abnormal situation based on the user's emotional state and respond immediately.
[0362] A "speech recognition device" is a device that converts speech into digital signals and has the function of detecting specific keywords or phrases in real time.
[0363] An "emotion recognition engine" is an engine equipped with an algorithm that analyzes a user's emotional state from voice data and environmental data and identifies specific emotions.
[0364] An "environmental information collecting device" is a device that acquires ambient environmental data such as temperature, light, and sound, and provides that data as basic information for analysis.
[0365] The "intelligent module function" is an artificial intelligence analysis function that comprehensively evaluates speech recognition, emotion recognition, and environmental information to determine the presence and urgency of abnormal situations.
[0366] The "communication function" refers to the function that performs digital communication to send notifications generated based on the analysis results to a designated terminal.
[0367] The system implementing this invention is comprised of a speech recognition device, an emotion recognition engine, a device for collecting environmental information, an intelligent module function, and a communication function. The server uses the speech recognition device to analyze voice commands from the user and detect specific emergency terms in real time. In this process, the voice data is captured as a digital signal using the SpeechRecognition library. Furthermore, the emotion recognition engine using TensorFlow analyzes the user's emotional state from the tone of the voice and ambient acoustics.
[0368] Meanwhile, the terminal acquires data such as temperature, light, and sound through a device that collects environmental information. The intelligent module function on the server comprehensively analyzes this data to determine whether there is an abnormal situation and how urgent it is. For example, if an unusually loud sound is detected at night and the user's emotions indicate strong anxiety or fear, the system will use that information to determine that a rapid response is necessary.
[0369] As a result, communication functions are used to send emergency notifications to relevant devices based on the results of analysis and sentiment evaluation. This allows appropriate administrators and stakeholders to quickly understand the situation and take necessary actions.
[0370] For example, if a user suddenly shouts "Help!!" at home one night, the system will instantly recognize the voice, and its emotion recognition engine will then identify any signs of anxiety. The server will analyze the situation and, if necessary, send a notification to the security administrator stating, "Unusual noise and strong anxiety in the middle of the night; immediate action required."
[0371] Examples of prompt messages include, "How does the system respond when an abnormal sound is detected?" and "What is the procedure when strong feelings of anxiety are detected?"
[0372] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0373] Step 1:
[0374] The terminal uses its built-in microphone to capture the user's voice and inputs it into a speech recognition device. The speech recognition device uses the SpeechRecognition library to convert the voice into text data and obtains an output that determines whether or not it contains emergency terms.
[0375] Step 2:
[0376] The device sends voice data and ambient sound environment data to the emotion recognition engine. The emotion recognition engine uses TensorFlow to analyze the tone and patterns of the voice and identify the user's emotional state. This process outputs an emotion classification result, such as fear or anxiety.
[0377] Step 3:
[0378] The server collects data acquired from devices that gather environmental information. Digital data such as temperature, light, and ambient noise levels are input into an analysis program to evaluate whether the environment is in an abnormal state. If an abnormal value is detected, the evaluation result is output.
[0379] Step 4:
[0380] The intelligent module function on the server comprehensively analyzes the speech recognition results and emotional state obtained from steps 1 and 2, and the environmental evaluation data from step 3. By correlating each input data, an analysis result is obtained that determines whether an abnormal situation exists and its level of urgency.
[0381] Step 5:
[0382] The server generates notifications via communication functions based on the analysis results of abnormal situations and urgency levels from the intelligent module. These notifications include the user's emotional state and environmental information, and are sent to the relevant administrator terminal as emergency notifications when a rapid response is required.
[0383] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0384] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0385] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0386] [Third Embodiment]
[0387] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0388] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0389] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0390] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0391] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0392] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0393] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0394] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0395] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0396] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0397] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0398] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0399] In an embodiment for carrying out the present invention, the system is configured as follows: This system includes a voice recognition sensor, a sensor for collecting environmental data, an artificial intelligence module, and communication means, all of which work together.
[0400] First, the device uses a voice recognition sensor to detect specific keywords, such as "Help!!". This automatically starts data collection in emergencies. In addition, sensors that collect environmental data continuously monitor surrounding environmental information such as temperature, movement, and light intensity. This provides a foundation for detecting abnormal situations that would not normally occur.
[0401] Next, the server receives the data collected from these sensors and analyzes it using an artificial intelligence module. This analysis involves pattern recognition for anomaly detection and comparison with historical data to determine any abnormalities. If an anomaly is detected, the server generates an emergency notification and transmits it to property managers and relevant authorities via communication channels.
[0402] Users receive notifications and take the necessary actions according to the instructions provided. In some cases, they can also contribute to system improvements by sending feedback to the server. Based on this feedback, the server adjusts system parameters to improve the accuracy of anomaly detection in the future.
[0403] For example, if a terminal detects the cry "Help!!", the audio data is immediately sent to a server for analysis. If it is determined that there is an emergency, such as someone having collapsed and being unable to move, the administrator is immediately notified. Through this series of processes, the system is structured to enable a rapid response to protect the safety of residents.
[0404] In this way, the present invention provides an embodiment that allows for a comprehensive understanding of the surrounding situation and immediate response in the event of an emergency.
[0405] The following describes the processing flow.
[0406] Step 1:
[0407] The device activates a voice recognition sensor to constantly monitor surrounding sounds and detects the keyword "Help!!". The voice recognition sensor analyzes specific voice patterns in real time and collects information when the keyword is detected.
[0408] Step 2:
[0409] The device activates sensors to collect environmental data in response to detection by the voice recognition sensor, recording information about the surrounding environment such as temperature, light intensity, and motion. This data is an important element for anomaly detection.
[0410] Step 3:
[0411] The device transmits collected audio and environmental data to the server. The data includes audio clips, various sensor information, and time information.
[0412] Step 4:
[0413] The server receives data sent from the terminal and stores it in the database. It then starts analyzing the received data using an artificial intelligence module.
[0414] Step 5:
[0415] The server uses an artificial intelligence module to analyze data, checking for the presence or absence of keywords in audio clips and for abnormal patterns in environmental data. If an anomaly is detected as a result of the analysis, an action is triggered.
[0416] Step 6:
[0417] When an anomaly is detected, the server determines what information to send in the notification and generates a notification message. This message includes the type of anomaly, the time it occurred, and the recommended action.
[0418] Step 7:
[0419] The server sends the generated notification message to the terminal of a pre-registered property manager or related organization. Notifications are sent via email or SMS using various communication methods.
[0420] Step 8:
[0421] Users receive notifications from the server on their devices and take necessary actions based on the content. For example, they might take emergency action to verify the site or ensure safety.
[0422] Step 9:
[0423] Users send feedback to the server regarding the results of their actions and areas for improvement. The server uses the received feedback to adjust system parameters and improve analysis accuracy.
[0424] (Example 1)
[0425] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0426] In recent years, there has been a growing demand for improved security and safety. However, conventional systems have a problem in that functions such as speech recognition, environmental data collection, and anomaly detection operate individually, making it difficult to respond quickly to abnormal situations. Furthermore, they lack mechanisms to improve system accuracy by utilizing user feedback. Therefore, there is a need to realize a system that can respond effectively and quickly to emergencies with a more integrated approach.
[0427] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0428] In this invention, the server includes voice recognition means for detecting predetermined keywords, means for collecting environmental information to monitor the surrounding conditions and detect anomalies, analysis means equipped with an intelligent module for processing the generated data, communication means for creating and transmitting warnings based on the analysis results, and means for receiving information from the user and adjusting system settings. This not only enables a rapid and accurate response to emergencies, but also allows for continuous improvement of the system's accuracy through user feedback.
[0429] A "designated keyword" is a specific word or phrase that has been set in advance and acts as a linguistic trigger that, when detected by speech recognition, causes a specific action.
[0430] "Voice recognition means" refers to a technology or device for sensing surrounding sounds and identifying specific keywords or phrases.
[0431] "Means of collecting environmental information" refers to devices and technologies that have the function of continuously monitoring the surrounding physical conditions, such as temperature, movement, and light intensity, and collecting them as data.
[0432] An "analysis means equipped with an intelligent module" is a computer system that includes algorithms or programs for performing learning and inference based on collected data, and for performing anomaly detection and pattern recognition.
[0433] "Communication methods" refer to technologies and devices that can send and receive data and warnings to remote locations, and typically utilize the internet or wireless communication.
[0434] "Means of receiving information from users and adjusting system settings" refers to technologies that allow a system to accept user feedback and requests and automatically modify or optimize its operation and parameters based on that information.
[0435] This invention provides an integrated system for detecting specific keywords, immediately recognizing abnormal environmental conditions, and taking appropriate action. Its configuration and operation are described below.
[0436] The device is equipped with a voice recognition sensor that constantly monitors the surrounding audio environment. This voice recognition sensor uses a high-performance microphone and voice analysis software to quickly identify situations requiring attention by detecting predetermined keywords, such as "Help!!".
[0437] Furthermore, the device has multiple environmental information gathering devices, including temperature, motion, and light sensors. These sensors continuously record the surrounding environmental conditions, providing a foundation for detecting signs of abnormal events or situations.
[0438] The server aggregates data sent from terminals and performs detailed analysis using an artificial intelligence module. This AI module has the ability to recognize anomalous patterns by comparing them with past data using machine learning algorithms. Based on the analysis results, the server can determine if an anomaly has occurred and immediately create an emergency notification.
[0439] Notifications generated through communication channels are immediately sent to property managers and relevant authorities. This enables a prompt response and protects the safety of residents.
[0440] Users take swift action based on received notifications. Depending on the situation, users can send feedback to the server. This feedback is used to optimize the system and improve the accuracy of anomaly detection in the future.
[0441] For example, if someone shouts "Help!!" in the hallway of an apartment building, the audio data is immediately sent to a server. If the audio analysis confirms that it is an urgent situation, a notification is sent to the administrator urging them to take immediate action.
[0442] A concrete example of a prompt for a generating AI model would be, "Please provide detailed information on the functions of sensors and communication systems for detecting and responding quickly to potential emergencies that may occur in real estate facilities." Based on this prompt, the generating AI would propose effective countermeasures appropriate to the situation.
[0443] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0444] Step 1:
[0445] The device constantly monitors ambient sounds using a voice recognition sensor. It receives ambient sound data as input and performs data processing to analyze whether it contains specific keywords. Specifically, the device performs frequency and phoneme analysis on the audio to detect the keyword "Help!!". Once the keyword is detected, it generates an alert and proceeds to the next step.
[0446] Step 2:
[0447] The device collects input data from temperature sensors, motion sensors, and light sensors. Based on this data, data processing is performed to identify unusual patterns. Specifically, real-time data from the sensors is integrated, and an algorithm is applied to detect anomalies. When an abnormal environmental change is detected, the aggregated data is sent to the server.
[0448] Step 3:
[0449] The server receives voice data and environmental data transmitted from the terminal. The input includes voice and environmental sensor data, and data calculations are performed using an artificial intelligence module to analyze this data. Specifically, anomaly patterns are recognized using a machine learning model and compared with past data to calculate the probability of an anomaly. If an anomaly is detected as a result of the analysis, the next step is to generate an emergency notification.
[0450] Step 4:
[0451] The server initiates a process to create notification content based on the analysis results. The input is the analyzed anomaly identification information, which is used to construct the notification message. Specifically, it automatically generates an emergency notification containing details of the situation for administrators or relevant authorities. Once the notification content is finalized, it is sent to the recipient via communication means.
[0452] Step 5:
[0453] The user receives a notification from the server. Based on the notification content, the user understands the situation and prepares to follow specific action instructions. Specifically, the user who checks the notification is expected to verify the situation on-site, take necessary actions, and send feedback to the server as needed, thereby contributing to improvements in the future.
[0454] (Application Example 1)
[0455] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0456] In recent years, public and commercial facilities have been required to respond quickly to emergencies. However, these facilities need to cover a wide area, making it difficult to continuously monitor all emergencies using only human resources. Therefore, there is a need for a system that can instantly analyze various environmental information, detect anomalies, and quickly notify relevant parties.
[0457] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0458] In this invention, the server includes means for a speech recognition detection device to sense specific instructional phrases, a sensor device for measuring numerical values of the external environment, and a machine learning module means for evaluating the acquired information. This makes it possible to quickly and accurately detect emergencies within public facilities and immediately notify relevant parties.
[0459] A "speech recognition detection device" is a device that detects specific instructional words as sounds and converts them into digital data.
[0460] A "sensor device" is a device that measures the physical characteristics and conditions of the external environment (e.g., temperature, light, motion) and collects them as data.
[0461] A "machine learning module" is a system consisting of algorithms and their implementations that use acquired information to detect patterns and anomalies and perform analysis.
[0462] A "communication device" is a device used to transmit notifications generated based on the analysis results to relevant terminals and equipment.
[0463] A "notification function" is a function that, when an emergency or abnormality is detected, sends appropriate notifications to pre-configured relevant parties or systems.
[0464] In an embodiment of this invention, the system is configured as follows: A server uses a speech recognition detection device to sense specific instructional phrases, which triggers a sensor device to measure numerical values of the external environment. The measured data is evaluated by a machine learning module to analyze whether or not there are any anomalies.
[0465] Specifically, the Google Speech Recognition API is used as the speech recognition detection device, converting specific instructional phrases into digital data when they are spoken. This resulting speech data, along with environmental data simultaneously acquired from sensor devices, is sent to a machine learning module for pattern recognition and anomaly detection.
[0466] Furthermore, the server generates a notification based on the analysis results via a communication device and quickly sends the notification to registered terminals. The notification function is activated, informing relevant parties of the emergency.
[0467] For example, if a voice shouting "Help!!" is detected inside a shopping mall, the footage and temperature data from the scene are analyzed, and a notification is sent to security personnel as needed. This notification includes detailed instructions for action, enabling a swift emergency response.
[0468] An example of a prompt to input into the generating AI model would be: "A voice suddenly shouting 'Help!!' was detected inside a shopping mall. The on-site video and temperature data are as follows... Based on this, please provide detailed instructions on how employees should respond."
[0469] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0470] Step 1:
[0471] Start voice recognition
[0472] The device initiates speech recognition using the Google Speech Recognition API. The input is ambient sound, and the device detects whether it contains the specific demonstrative phrase "Help!!". If detected, the audio data is sent to the server as digital data.
[0473] Step 2:
[0474] Environmental data collection
[0475] The server measures external environmental data from sensor devices in parallel with speech recognition. These measurements include temperature sensors and motion detection sensors, and the collected data includes temperature, motion, and light intensity. The input here is the sensor values, and the output is the aggregated environmental data.
[0476] Step 3:
[0477] Data Analysis
[0478] The server sends the acquired audio and environmental data to a machine learning module. Based on the input data, the module analyzes the data to determine if it is an anomaly by comparing it to past patterns. The output is the result of the anomaly determination. Specifically, if there is a possibility of an anomaly, the module also calculates the confidence level of that determination.
[0479] Step 4:
[0480] Notification generation
[0481] If an anomaly is detected, the server generates a notification via the communication device. The input is the result of the anomaly analysis, and the output is an emergency notification message. Specifically, the message includes a description of the situation on site and recommended actions.
[0482] Step 5:
[0483] Send notification
[0484] The server sends generated notifications to registered devices. The input is the notification message, and the output is the notification displayed on the device. This allows users to confirm the details of an emergency and take prompt action. The notification provides specific action guidelines, ensuring that appropriate instructions are conveyed to employees and relevant parties.
[0485] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0486] In an embodiment for carrying out the present invention, the system comprises a voice recognition sensor, an environmental data collection sensor, an artificial intelligence module, a communication means, and an emotion recognition engine.
[0487] First, the device uses a voice recognition sensor to continuously monitor for emergency keywords such as "Help!!". In addition, an emotion recognition engine analyzes the user's voice and surrounding sounds to determine their emotional state. This analysis makes it possible to understand whether the user is stressed or calm.
[0488] Next, the device sends the collected data to the server. The server processes the received data using an artificial intelligence module. This process combines voice and environmental data for detailed analysis to understand the user's situation. Emotional state data provided by the emotion recognition engine is also used as part of the analysis. This allows for the detection of abnormal conditions that require immediate attention.
[0489] When an anomaly is detected, the server generates a notification depending on the situation. At this time, emotional information obtained from the emotion recognition engine is used to customize the notification message, resulting in a more appropriate notification. For example, if the user is experiencing extreme anxiety or fear, this will be reflected in the notification, prompting the administrator receiving the notification to take prompt action.
[0490] Users receive notifications from the server and take action based on the situation on-site. The notification recipients can immediately understand the urgency and severity of the situation through messages that include emotional information. Feedback received after the response is sent back to the server, which is used to help the system learn and improve.
[0491] For example, if the emergency keyword "Help!!" is recognized, and the emotion recognition engine determines from the tone of voice that the user is in a state of extreme fear, a notification is generated more quickly than usual, prompting immediate action. Based on this information, the server sends an emergency response alert to the property manager.
[0492] In this way, the system, which incorporates an emotion recognition engine, has a structure that enables more appropriate and rapid responses by taking the user's emotional state into consideration during decision-making.
[0493] The following describes the processing flow.
[0494] Step 1:
[0495] The device constantly monitors ambient sounds using a voice recognition sensor to detect whether "Help!!" or other emergency keywords are spoken. Additionally, an emotion engine analyzes the user's voice tone to recognize emotional states such as stress or anxiety.
[0496] Step 2:
[0497] The device transmits identified voice data and emotional state data to the server. Simultaneously, data such as temperature, light intensity, and motion collected by environmental sensors are also transferred to the server. This provides comprehensive information about the user's state and living environment.
[0498] Step 3:
[0499] The server receives data sent from the terminal and stores it in a database. This data is analyzed using an artificial intelligence module to check for abnormalities in voice keywords and emotional states.
[0500] Step 4:
[0501] The server determines whether an anomaly has been detected based on the analysis results, including the output of the emotion engine. If the emotional state is particularly unstable, it is more likely to be treated as an anomaly.
[0502] Step 5:
[0503] If an anomaly is detected, the server generates a notification message tailored to the nature of the anomaly. This message is customized based on the user's emotional state, with its wording adjusted based on the level of urgency and the corresponding emotion.
[0504] Step 6:
[0505] The server sends the generated notification message to the terminal of the registered property manager or relevant organization. This is done using standard communication methods such as email or SMS.
[0506] Step 7:
[0507] Users receive notifications from the server and take prompt action based on their content. These notifications include the cause of the anomaly and the user's estimated emotional state, allowing for a more accurate understanding of the situation.
[0508] Step 8:
[0509] After completing the task, users send feedback to the server via their device regarding the results and any discoveries they made. This feedback is used as training data for the system and contributes to future system improvements.
[0510] (Example 2)
[0511] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0512] Conventional speech recognition systems primarily rely on keyword detection for anomaly detection, making it difficult to provide rapid and appropriate responses that take into account the user's emotional state. Furthermore, they lacked mechanisms to comprehensively analyze environmental information to determine urgency and provide appropriate notifications. Additionally, they lacked effective means of utilizing user feedback for system improvement.
[0513] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0514] In this invention, the server includes means for a speech recognition device to monitor specified keywords, a collection device for collecting environmental information and detecting anomalies, a machine learning module for evaluating the analyzed data in detail, and an emotion recognition engine for determining and utilizing emotional states in analysis. This makes it possible to comprehensively analyze the user's emotional state and environmental information and generate quick and appropriate notifications. Furthermore, user feedback can be used to optimize the system.
[0515] A "voice recognition device" is a device that continuously monitors surrounding sounds and has the function of detecting specific keywords.
[0516] A "collection device" is a device used to collect information about the surrounding environment and detect anomalies.
[0517] A "machine learning module" is a system that includes algorithms and models for performing detailed analysis on collected data and evaluating the information.
[0518] A "communication device" is a device that generates notifications based on analysis results and transmits them externally.
[0519] An "emotion recognition engine" refers to a technology that identifies emotional states from voice data and uses that information for analysis.
[0520] This system includes a speech recognition device, data collection device, machine learning module, emotion recognition engine, and communication device for the user. The terminal uses the speech recognition device to acquire voice data from the user's surroundings. This device has the function of continuously monitoring specific keywords and detecting keywords for emergencies.
[0521] Furthermore, the terminal is equipped with a data collection device for collecting environmental data. This device is intended to collect information about the user's surroundings and detect anomalies. The collected data is transmitted from the terminal to a server.
[0522] The server uses a machine learning module to analyze voice data and environmental information in detail. This module comprehensively evaluates the data based on a generative AI model to determine the user's situation. It also uses an emotion recognition engine to identify the user's emotional state from their voice and incorporates this into the analysis results.
[0523] For example, if the keyword "Help!!" is detected and strong fear is perceived from the user's voice, this information will be included in the analysis results. Based on the analysis results, the server will use a communication device to automatically generate an appropriate notification and send it to the administrator.
[0524] Upon receiving a notification, the user takes action based on the content of the notification, according to the situation at hand. An example of a prompt message, used when analyzing a specific emotional state, is, "Analyze the emotional state of this audio sample and determine whether it is an emergency." This allows the system to ensure the user's safety in real time and enable a rapid response.
[0525] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0526] Step 1:
[0527] The device acquires surrounding audio in real time using a speech recognition device. The input is raw audio data from the surroundings. The speech recognition device processes this data and detects a specified keyword, such as "Help!!". The output is a flag indicating that the keyword was detected. Specifically, when audio is input, a high-speed speech analysis algorithm is applied to determine the presence or absence of the keyword.
[0528] Step 2:
[0529] The device uses an emotion recognition engine to perform emotion analysis on the audio data acquired in Step 1. The input is audio data, and the emotion recognition engine analyzes the tone, speed, and volume of the voice to determine the user's emotional state. The output is data indicating the emotional state. For example, if the voice contains clear signs of anxiety or fear, that emotion will be detected.
[0530] Step 3:
[0531] The terminal uses a data collection device to gather information about the surrounding environment. The input is environmental data from sensors. The data collection device acquires information such as temperature, humidity, and brightness, and checks for any abnormal values. The output is a set of current environmental information. For example, if there is an abnormal temperature rise, that information is recorded.
[0532] Step 4:
[0533] The terminal sends the data obtained in steps 1, 2, and 3 to the server. The input consists of voice keyword flags, emotional state data, and environmental information. This data is sent to the server via the internet. The output is a signal confirming that this data has arrived at the server. A secure protocol is used to ensure the security of the data.
[0534] Step 5:
[0535] The server uses a machine learning module to analyze data sent from the terminal in detail. Inputs include voice keyword flags, emotional state data, and environmental information. The machine learning module uses a generative AI model to comprehensively evaluate the data and determine the user's situation. The output is an evaluation result indicating the user's situation. For example, if an emergency is determined, that information will be generated.
[0536] Step 6:
[0537] The server generates notifications based on the analysis results and sends the appropriate notifications to the administrator using a communication device. The input is the server's evaluation result. A dedicated template is used to generate the content of the notification. The output is the generated notification message. For example, a notification indicating an emergency is sent to the administrator.
[0538] Step 7:
[0539] The user receives notifications from the server and takes appropriate action on-site based on the notification content. The input is the notification message sent from the server. The user understands its content and takes specific action. The output is the user's action. If feedback information is obtained, it is also sent to the server.
[0540] (Application Example 2)
[0541] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0542] Current security systems lack sufficient emergency response capabilities that take into account the user's emotional state, making it difficult to immediately assess the severity and urgency of an abnormal situation. Therefore, appropriate and rapid responses tailored to the situation are required.
[0543] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0544] In this invention, the server includes means for a speech recognition device to detect predetermined emergency terms, means for analyzing the user's emotional state using an emotion recognition engine, and a device for collecting environmental information. This makes it possible to determine the urgency of an abnormal situation based on the user's emotional state and respond immediately.
[0545] A "speech recognition device" is a device that converts speech into digital signals and has the function of detecting specific keywords or phrases in real time.
[0546] An "emotion recognition engine" is an engine equipped with an algorithm that analyzes a user's emotional state from voice data and environmental data and identifies specific emotions.
[0547] An "environmental information collecting device" is a device that acquires ambient environmental data such as temperature, light, and sound, and provides that data as basic information for analysis.
[0548] The "intelligent module function" is an artificial intelligence analysis function that comprehensively evaluates speech recognition, emotion recognition, and environmental information to determine the presence and urgency of abnormal situations.
[0549] The "communication function" refers to the function that performs digital communication to send notifications generated based on the analysis results to a designated terminal.
[0550] The system implementing this invention is comprised of a speech recognition device, an emotion recognition engine, a device for collecting environmental information, an intelligent module function, and a communication function. The server uses the speech recognition device to analyze voice commands from the user and detect specific emergency terms in real time. In this process, the voice data is captured as a digital signal using the SpeechRecognition library. Furthermore, the emotion recognition engine using TensorFlow analyzes the user's emotional state from the tone of the voice and ambient acoustics.
[0551] Meanwhile, the terminal acquires data such as temperature, light, and sound through a device that collects environmental information. The intelligent module function on the server comprehensively analyzes this data to determine whether there is an abnormal situation and how urgent it is. For example, if an unusually loud sound is detected at night and the user's emotions indicate strong anxiety or fear, the system will use that information to determine that a rapid response is necessary.
[0552] As a result, communication functions are used to send emergency notifications to relevant devices based on the results of analysis and sentiment evaluation. This allows appropriate administrators and stakeholders to quickly understand the situation and take necessary actions.
[0553] For example, if a user suddenly shouts "Help!!" at home one night, the system will instantly recognize the voice, and its emotion recognition engine will then identify any signs of anxiety. The server will analyze the situation and, if necessary, send a notification to the security administrator stating, "Unusual noise and strong anxiety in the middle of the night; immediate action required."
[0554] Examples of prompt messages include, "How does the system respond when an abnormal sound is detected?" and "What is the procedure when strong feelings of anxiety are detected?"
[0555] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0556] Step 1:
[0557] The terminal uses its built-in microphone to capture the user's voice and inputs it into a speech recognition device. The speech recognition device uses the SpeechRecognition library to convert the voice into text data and obtains an output that determines whether or not it contains emergency terms.
[0558] Step 2:
[0559] The device sends voice data and ambient sound environment data to the emotion recognition engine. The emotion recognition engine uses TensorFlow to analyze the tone and patterns of the voice and identify the user's emotional state. This process outputs an emotion classification result, such as fear or anxiety.
[0560] Step 3:
[0561] The server collects data acquired from devices that gather environmental information. Digital data such as temperature, light, and ambient noise levels are input into an analysis program to evaluate whether the environment is in an abnormal state. If an abnormal value is detected, the evaluation result is output.
[0562] Step 4:
[0563] The intelligent module function on the server comprehensively analyzes the speech recognition results and emotional state obtained from steps 1 and 2, and the environmental evaluation data from step 3. By correlating each input data, an analysis result is obtained that determines whether an abnormal situation exists and its level of urgency.
[0564] Step 5:
[0565] The server generates notifications via communication functions based on the analysis results of abnormal situations and urgency levels from the intelligent module. These notifications include the user's emotional state and environmental information, and are sent to the relevant administrator terminal as emergency notifications when a rapid response is required.
[0566] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0567] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0568] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0569] [Fourth Embodiment]
[0570] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0571] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0572] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0573] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0574] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0575] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0576] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0577] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0578] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0579] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0580] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0581] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0582] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0583] In an embodiment for carrying out the present invention, the system is configured as follows: This system includes a voice recognition sensor, a sensor for collecting environmental data, an artificial intelligence module, and communication means, all of which work together.
[0584] First, the device uses a voice recognition sensor to detect specific keywords, such as "Help!!". This automatically starts data collection in emergencies. In addition, sensors that collect environmental data continuously monitor surrounding environmental information such as temperature, movement, and light intensity. This provides a foundation for detecting abnormal situations that would not normally occur.
[0585] Next, the server receives the data collected from these sensors and analyzes it using an artificial intelligence module. This analysis involves pattern recognition for anomaly detection and comparison with historical data to determine any abnormalities. If an anomaly is detected, the server generates an emergency notification and transmits it to property managers and relevant authorities via communication channels.
[0586] Users receive notifications and take the necessary actions according to the instructions provided. In some cases, they can also contribute to system improvements by sending feedback to the server. Based on this feedback, the server adjusts system parameters to improve the accuracy of anomaly detection in the future.
[0587] For example, if a terminal detects the cry "Help!!", the audio data is immediately sent to a server for analysis. If it is determined that there is an emergency, such as someone having collapsed and being unable to move, the administrator is immediately notified. Through this series of processes, the system is structured to enable a rapid response to protect the safety of residents.
[0588] In this way, the present invention provides an embodiment that allows for a comprehensive understanding of the surrounding situation and immediate response in the event of an emergency.
[0589] The following describes the processing flow.
[0590] Step 1:
[0591] The device activates a voice recognition sensor to constantly monitor surrounding sounds and detects the keyword "Help!!". The voice recognition sensor analyzes specific voice patterns in real time and collects information when the keyword is detected.
[0592] Step 2:
[0593] The device activates sensors to collect environmental data in response to detection by the voice recognition sensor, recording information about the surrounding environment such as temperature, light intensity, and motion. This data is an important element for anomaly detection.
[0594] Step 3:
[0595] The device transmits collected audio and environmental data to the server. The data includes audio clips, various sensor information, and time information.
[0596] Step 4:
[0597] The server receives data sent from the terminal and stores it in the database. It then starts analyzing the received data using an artificial intelligence module.
[0598] Step 5:
[0599] The server uses an artificial intelligence module to analyze data, checking for the presence or absence of keywords in audio clips and for abnormal patterns in environmental data. If an anomaly is detected as a result of the analysis, an action is triggered.
[0600] Step 6:
[0601] When an anomaly is detected, the server determines what information to send in the notification and generates a notification message. This message includes the type of anomaly, the time it occurred, and the recommended action.
[0602] Step 7:
[0603] The server sends the generated notification message to the terminal of a pre-registered property manager or related organization. Notifications are sent via email or SMS using various communication methods.
[0604] Step 8:
[0605] Users receive notifications from the server on their devices and take necessary actions based on the content. For example, they might take emergency action to verify the site or ensure safety.
[0606] Step 9:
[0607] Users send feedback to the server regarding the results of their actions and areas for improvement. The server uses the received feedback to adjust system parameters and improve analysis accuracy.
[0608] (Example 1)
[0609] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0610] In recent years, there has been a growing demand for improved security and safety. However, conventional systems have a problem in that functions such as speech recognition, environmental data collection, and anomaly detection operate individually, making it difficult to respond quickly to abnormal situations. Furthermore, they lack mechanisms to improve system accuracy by utilizing user feedback. Therefore, there is a need to realize a system that can respond effectively and quickly to emergencies with a more integrated approach.
[0611] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0612] In this invention, the server includes voice recognition means for detecting predetermined keywords, means for collecting environmental information to monitor the surrounding conditions and detect anomalies, analysis means equipped with an intelligent module for processing the generated data, communication means for creating and transmitting warnings based on the analysis results, and means for receiving information from the user and adjusting system settings. This not only enables a rapid and accurate response to emergencies, but also allows for continuous improvement of the system's accuracy through user feedback.
[0613] A "designated keyword" is a specific word or phrase that has been set in advance and acts as a linguistic trigger that, when detected by speech recognition, causes a specific action.
[0614] "Voice recognition means" refers to a technology or device for sensing surrounding sounds and identifying specific keywords or phrases.
[0615] "Means of collecting environmental information" refers to devices and technologies that have the function of continuously monitoring the surrounding physical conditions, such as temperature, movement, and light intensity, and collecting them as data.
[0616] An "analysis means equipped with an intelligent module" is a computer system that includes algorithms or programs for performing learning and inference based on collected data, and for performing anomaly detection and pattern recognition.
[0617] "Communication methods" refer to technologies and devices that can send and receive data and warnings to remote locations, and typically utilize the internet or wireless communication.
[0618] "Means of receiving information from users and adjusting system settings" refers to technologies that allow a system to accept user feedback and requests and automatically modify or optimize its operation and parameters based on that information.
[0619] This invention provides an integrated system for detecting specific keywords, immediately recognizing abnormal environmental conditions, and taking appropriate action. Its configuration and operation are described below.
[0620] The device is equipped with a voice recognition sensor that constantly monitors the surrounding audio environment. This voice recognition sensor uses a high-performance microphone and voice analysis software to quickly identify situations requiring attention by detecting predetermined keywords, such as "Help!!".
[0621] Furthermore, the device has multiple environmental information gathering devices, including temperature, motion, and light sensors. These sensors continuously record the surrounding environmental conditions, providing a foundation for detecting signs of abnormal events or situations.
[0622] The server aggregates data sent from terminals and performs detailed analysis using an artificial intelligence module. This AI module has the ability to recognize anomalous patterns by comparing them with past data using machine learning algorithms. Based on the analysis results, the server can determine if an anomaly has occurred and immediately create an emergency notification.
[0623] Notifications generated through communication channels are immediately sent to property managers and relevant authorities. This enables a prompt response and protects the safety of residents.
[0624] Users take swift action based on received notifications. Depending on the situation, users can send feedback to the server. This feedback is used to optimize the system and improve the accuracy of anomaly detection in the future.
[0625] For example, if someone shouts "Help!!" in the hallway of an apartment building, the audio data is immediately sent to a server. If the audio analysis confirms that it is an urgent situation, a notification is sent to the administrator urging them to take immediate action.
[0626] A concrete example of a prompt for a generating AI model would be, "Please provide detailed information on the functions of sensors and communication systems for detecting and responding quickly to potential emergencies that may occur in real estate facilities." Based on this prompt, the generating AI would propose effective countermeasures appropriate to the situation.
[0627] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0628] Step 1:
[0629] The device constantly monitors ambient sounds using a voice recognition sensor. It receives ambient sound data as input and performs data processing to analyze whether it contains specific keywords. Specifically, the device performs frequency and phoneme analysis on the audio to detect the keyword "Help!!". Once the keyword is detected, it generates an alert and proceeds to the next step.
[0630] Step 2:
[0631] The device collects input data from temperature sensors, motion sensors, and light sensors. Based on this data, data processing is performed to identify unusual patterns. Specifically, real-time data from the sensors is integrated, and an algorithm is applied to detect anomalies. When an abnormal environmental change is detected, the aggregated data is sent to the server.
[0632] Step 3:
[0633] The server receives voice data and environmental data transmitted from the terminal. The input includes voice and environmental sensor data, and data calculations are performed using an artificial intelligence module to analyze this data. Specifically, anomaly patterns are recognized using a machine learning model and compared with past data to calculate the probability of an anomaly. If an anomaly is detected as a result of the analysis, the next step is to generate an emergency notification.
[0634] Step 4:
[0635] The server initiates a process to create notification content based on the analysis results. The input is the analyzed anomaly identification information, which is used to construct the notification message. Specifically, it automatically generates an emergency notification containing details of the situation for administrators or relevant authorities. Once the notification content is finalized, it is sent to the recipient via communication means.
[0636] Step 5:
[0637] The user receives a notification from the server. Based on the notification content, the user understands the situation and prepares to follow specific action instructions. Specifically, the user who checks the notification is expected to verify the situation on-site, take necessary actions, and send feedback to the server as needed, thereby contributing to improvements in the future.
[0638] (Application Example 1)
[0639] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0640] In recent years, public and commercial facilities have been required to respond quickly to emergencies. However, these facilities need to cover a wide area, making it difficult to continuously monitor all emergencies using only human resources. Therefore, there is a need for a system that can instantly analyze various environmental information, detect anomalies, and quickly notify relevant parties.
[0641] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0642] In this invention, the server includes means for a speech recognition detection device to sense specific instructional phrases, a sensor device for measuring numerical values of the external environment, and a machine learning module means for evaluating the acquired information. This makes it possible to quickly and accurately detect emergencies within public facilities and immediately notify relevant parties.
[0643] A "speech recognition detection device" is a device that detects specific instructional words as sounds and converts them into digital data.
[0644] A "sensor device" is a device that measures the physical characteristics and conditions of the external environment (e.g., temperature, light, motion) and collects them as data.
[0645] A "machine learning module" is a system consisting of algorithms and their implementations that use acquired information to detect patterns and anomalies and perform analysis.
[0646] A "communication device" is a device used to transmit notifications generated based on the analysis results to relevant terminals and equipment.
[0647] A "notification function" is a function that, when an emergency or abnormality is detected, sends appropriate notifications to pre-configured relevant parties or systems.
[0648] In an embodiment of this invention, the system is configured as follows: A server uses a speech recognition detection device to sense specific instructional phrases, which triggers a sensor device to measure numerical values of the external environment. The measured data is evaluated by a machine learning module to analyze whether or not there are any anomalies.
[0649] Specifically, the Google Speech Recognition API is used as the speech recognition detection device, converting specific instructional phrases into digital data when they are spoken. This resulting speech data, along with environmental data simultaneously acquired from sensor devices, is sent to a machine learning module for pattern recognition and anomaly detection.
[0650] Furthermore, the server generates a notification based on the analysis results via a communication device and quickly sends the notification to registered terminals. The notification function is activated, informing relevant parties of the emergency.
[0651] For example, if a voice shouting "Help!!" is detected inside a shopping mall, the footage and temperature data from the scene are analyzed, and a notification is sent to security personnel as needed. This notification includes detailed instructions for action, enabling a swift emergency response.
[0652] An example of a prompt to input into the generating AI model would be: "A voice suddenly shouting 'Help!!' was detected inside a shopping mall. The on-site video and temperature data are as follows... Based on this, please provide detailed instructions on how employees should respond."
[0653] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0654] Step 1:
[0655] Start voice recognition
[0656] The device initiates speech recognition using the Google Speech Recognition API. The input is ambient sound, and the device detects whether it contains the specific demonstrative phrase "Help!!". If detected, the audio data is sent to the server as digital data.
[0657] Step 2:
[0658] Environmental data collection
[0659] The server measures external environmental data from sensor devices in parallel with speech recognition. These measurements include temperature sensors and motion detection sensors, and the collected data includes temperature, motion, and light intensity. The input here is the sensor values, and the output is the aggregated environmental data.
[0660] Step 3:
[0661] Data Analysis
[0662] The server sends the acquired audio and environmental data to a machine learning module. Based on the input data, the module analyzes the data to determine if it is an anomaly by comparing it to past patterns. The output is the result of the anomaly determination. Specifically, if there is a possibility of an anomaly, the module also calculates the confidence level of that determination.
[0663] Step 4:
[0664] Notification generation
[0665] If an anomaly is detected, the server generates a notification via the communication device. The input is the result of the anomaly analysis, and the output is an emergency notification message. Specifically, the message includes a description of the situation on site and recommended actions.
[0666] Step 5:
[0667] Send notification
[0668] The server sends generated notifications to registered devices. The input is the notification message, and the output is the notification displayed on the device. This allows users to confirm the details of an emergency and take prompt action. The notification provides specific action guidelines, ensuring that appropriate instructions are conveyed to employees and relevant parties.
[0669] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0670] In an embodiment for carrying out the present invention, the system comprises a voice recognition sensor, an environmental data collection sensor, an artificial intelligence module, a communication means, and an emotion recognition engine.
[0671] First, the device uses a voice recognition sensor to continuously monitor for emergency keywords such as "Help!!". In addition, an emotion recognition engine analyzes the user's voice and surrounding sounds to determine their emotional state. This analysis makes it possible to understand whether the user is stressed or calm.
[0672] Next, the device sends the collected data to the server. The server processes the received data using an artificial intelligence module. This process combines voice and environmental data for detailed analysis to understand the user's situation. Emotional state data provided by the emotion recognition engine is also used as part of the analysis. This allows for the detection of abnormal conditions that require immediate attention.
[0673] When an anomaly is detected, the server generates a notification depending on the situation. At this time, emotional information obtained from the emotion recognition engine is used to customize the notification message, resulting in a more appropriate notification. For example, if the user is experiencing extreme anxiety or fear, this will be reflected in the notification, prompting the administrator receiving the notification to take prompt action.
[0674] Users receive notifications from the server and take action based on the situation on-site. The notification recipients can immediately understand the urgency and severity of the situation through messages that include emotional information. Feedback received after the response is sent back to the server, which is used to help the system learn and improve.
[0675] For example, if the emergency keyword "Help!!" is recognized, and the emotion recognition engine determines from the tone of voice that the user is in a state of extreme fear, a notification is generated more quickly than usual, prompting immediate action. Based on this information, the server sends an emergency response alert to the property manager.
[0676] In this way, the system, which incorporates an emotion recognition engine, has a structure that enables more appropriate and rapid responses by taking the user's emotional state into consideration during decision-making.
[0677] The following describes the processing flow.
[0678] Step 1:
[0679] The device constantly monitors ambient sounds using a voice recognition sensor to detect whether "Help!!" or other emergency keywords are spoken. Additionally, an emotion engine analyzes the user's voice tone to recognize emotional states such as stress or anxiety.
[0680] Step 2:
[0681] The device transmits identified voice data and emotional state data to the server. Simultaneously, data such as temperature, light intensity, and motion collected by environmental sensors are also transferred to the server. This provides comprehensive information about the user's state and living environment.
[0682] Step 3:
[0683] The server receives data sent from the terminal and stores it in a database. This data is analyzed using an artificial intelligence module to check for abnormalities in voice keywords and emotional states.
[0684] Step 4:
[0685] The server determines whether an anomaly has been detected based on the analysis results, including the output of the emotion engine. If the emotional state is particularly unstable, it is more likely to be treated as an anomaly.
[0686] Step 5:
[0687] If an anomaly is detected, the server generates a notification message tailored to the nature of the anomaly. This message is customized based on the user's emotional state, with its wording adjusted based on the level of urgency and the corresponding emotion.
[0688] Step 6:
[0689] The server sends the generated notification message to the terminal of the registered property manager or relevant organization. This is done using standard communication methods such as email or SMS.
[0690] Step 7:
[0691] Users receive notifications from the server and take prompt action based on their content. These notifications include the cause of the anomaly and the user's estimated emotional state, allowing for a more accurate understanding of the situation.
[0692] Step 8:
[0693] After completing the task, users send feedback to the server via their device regarding the results and any discoveries they made. This feedback is used as training data for the system and contributes to future system improvements.
[0694] (Example 2)
[0695] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0696] Conventional speech recognition systems primarily rely on keyword detection for anomaly detection, making it difficult to provide rapid and appropriate responses that take into account the user's emotional state. Furthermore, they lacked mechanisms to comprehensively analyze environmental information to determine urgency and provide appropriate notifications. Additionally, they lacked effective means of utilizing user feedback for system improvement.
[0697] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0698] In this invention, the server includes means for a speech recognition device to monitor specified keywords, a collection device for collecting environmental information and detecting anomalies, a machine learning module for evaluating the analyzed data in detail, and an emotion recognition engine for determining and utilizing emotional states in analysis. This makes it possible to comprehensively analyze the user's emotional state and environmental information and generate quick and appropriate notifications. Furthermore, user feedback can be used to optimize the system.
[0699] A "voice recognition device" is a device that continuously monitors surrounding sounds and has the function of detecting specific keywords.
[0700] A "collection device" is a device used to collect information about the surrounding environment and detect anomalies.
[0701] A "machine learning module" is a system that includes algorithms and models for performing detailed analysis on collected data and evaluating the information.
[0702] A "communication device" is a device that generates notifications based on analysis results and transmits them externally.
[0703] An "emotion recognition engine" refers to a technology that identifies emotional states from voice data and uses that information for analysis.
[0704] This system includes a speech recognition device, data collection device, machine learning module, emotion recognition engine, and communication device for the user. The terminal uses the speech recognition device to acquire voice data from the user's surroundings. This device has the function of continuously monitoring specific keywords and detecting keywords for emergencies.
[0705] Furthermore, the terminal is equipped with a data collection device for collecting environmental data. This device is intended to collect information about the user's surroundings and detect anomalies. The collected data is transmitted from the terminal to a server.
[0706] The server uses a machine learning module to analyze voice data and environmental information in detail. This module comprehensively evaluates the data based on a generative AI model to determine the user's situation. It also uses an emotion recognition engine to identify the user's emotional state from their voice and incorporates this into the analysis results.
[0707] For example, if the keyword "Help!!" is detected and strong fear is perceived from the user's voice, this information will be included in the analysis results. Based on the analysis results, the server will use a communication device to automatically generate an appropriate notification and send it to the administrator.
[0708] Upon receiving a notification, the user takes action based on the content of the notification, according to the situation at hand. An example of a prompt message, used when analyzing a specific emotional state, is, "Analyze the emotional state of this audio sample and determine whether it is an emergency." This allows the system to ensure the user's safety in real time and enable a rapid response.
[0709] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0710] Step 1:
[0711] The device acquires surrounding audio in real time using a speech recognition device. The input is raw audio data from the surroundings. The speech recognition device processes this data and detects a specified keyword, such as "Help!!". The output is a flag indicating that the keyword was detected. Specifically, when audio is input, a high-speed speech analysis algorithm is applied to determine the presence or absence of the keyword.
[0712] Step 2:
[0713] The device uses an emotion recognition engine to perform emotion analysis on the audio data acquired in Step 1. The input is audio data, and the emotion recognition engine analyzes the tone, speed, and volume of the voice to determine the user's emotional state. The output is data indicating the emotional state. For example, if the voice contains clear signs of anxiety or fear, that emotion will be detected.
[0714] Step 3:
[0715] The terminal uses a data collection device to gather information about the surrounding environment. The input is environmental data from sensors. The data collection device acquires information such as temperature, humidity, and brightness, and checks for any abnormal values. The output is a set of current environmental information. For example, if there is an abnormal temperature rise, that information is recorded.
[0716] Step 4:
[0717] The terminal sends the data obtained in steps 1, 2, and 3 to the server. The input consists of voice keyword flags, emotional state data, and environmental information. This data is sent to the server via the internet. The output is a signal confirming that this data has arrived at the server. A secure protocol is used to ensure the security of the data.
[0718] Step 5:
[0719] The server uses a machine learning module to analyze data sent from the terminal in detail. Inputs include voice keyword flags, emotional state data, and environmental information. The machine learning module uses a generative AI model to comprehensively evaluate the data and determine the user's situation. The output is an evaluation result indicating the user's situation. For example, if an emergency is determined, that information will be generated.
[0720] Step 6:
[0721] The server generates notifications based on the analysis results and sends the appropriate notifications to the administrator using a communication device. The input is the server's evaluation result. A dedicated template is used to generate the content of the notification. The output is the generated notification message. For example, a notification indicating an emergency is sent to the administrator.
[0722] Step 7:
[0723] The user receives notifications from the server and takes appropriate action on-site based on the notification content. The input is the notification message sent from the server. The user understands its content and takes specific action. The output is the user's action. If feedback information is obtained, it is also sent to the server.
[0724] (Application Example 2)
[0725] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0726] Current security systems lack sufficient emergency response capabilities that take into account the user's emotional state, making it difficult to immediately assess the severity and urgency of an abnormal situation. Therefore, appropriate and rapid responses tailored to the situation are required.
[0727] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0728] In this invention, the server includes means for a speech recognition device to detect predetermined emergency terms, means for analyzing the user's emotional state using an emotion recognition engine, and a device for collecting environmental information. This makes it possible to determine the urgency of an abnormal situation based on the user's emotional state and respond immediately.
[0729] A "speech recognition device" is a device that converts speech into digital signals and has the function of detecting specific keywords or phrases in real time.
[0730] An "emotion recognition engine" is an engine equipped with an algorithm that analyzes a user's emotional state from voice data and environmental data and identifies specific emotions.
[0731] An "environmental information collecting device" is a device that acquires ambient environmental data such as temperature, light, and sound, and provides that data as basic information for analysis.
[0732] The "intelligent module function" is an artificial intelligence analysis function that comprehensively evaluates speech recognition, emotion recognition, and environmental information to determine the presence and urgency of abnormal situations.
[0733] The "communication function" refers to the function that performs digital communication to send notifications generated based on the analysis results to a designated terminal.
[0734] The system implementing this invention is comprised of a speech recognition device, an emotion recognition engine, a device for collecting environmental information, an intelligent module function, and a communication function. The server uses the speech recognition device to analyze voice commands from the user and detect specific emergency terms in real time. In this process, the voice data is captured as a digital signal using the SpeechRecognition library. Furthermore, the emotion recognition engine using TensorFlow analyzes the user's emotional state from the tone of the voice and ambient acoustics.
[0735] Meanwhile, the terminal acquires data such as temperature, light, and sound through a device that collects environmental information. The intelligent module function on the server comprehensively analyzes this data to determine whether there is an abnormal situation and how urgent it is. For example, if an unusually loud sound is detected at night and the user's emotions indicate strong anxiety or fear, the system will use that information to determine that a rapid response is necessary.
[0736] As a result, communication functions are used to send emergency notifications to relevant devices based on the results of analysis and sentiment evaluation. This allows appropriate administrators and stakeholders to quickly understand the situation and take necessary actions.
[0737] For example, if a user suddenly shouts "Help!!" at home one night, the system will instantly recognize the voice, and its emotion recognition engine will then identify any signs of anxiety. The server will analyze the situation and, if necessary, send a notification to the security administrator stating, "Unusual noise and strong anxiety in the middle of the night; immediate action required."
[0738] Examples of prompt messages include, "How does the system respond when an abnormal sound is detected?" and "What is the procedure when strong feelings of anxiety are detected?"
[0739] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0740] Step 1:
[0741] The terminal uses its built-in microphone to capture the user's voice and inputs it into a speech recognition device. The speech recognition device uses the SpeechRecognition library to convert the voice into text data and obtains an output that determines whether or not it contains emergency terms.
[0742] Step 2:
[0743] The device sends voice data and ambient sound environment data to the emotion recognition engine. The emotion recognition engine uses TensorFlow to analyze the tone and patterns of the voice and identify the user's emotional state. This process outputs an emotion classification result, such as fear or anxiety.
[0744] Step 3:
[0745] The server collects data acquired from devices that gather environmental information. Digital data such as temperature, light, and ambient noise levels are input into an analysis program to evaluate whether the environment is in an abnormal state. If an abnormal value is detected, the evaluation result is output.
[0746] Step 4:
[0747] The intelligent module function on the server comprehensively analyzes the speech recognition results and emotional state obtained from steps 1 and 2, and the environmental evaluation data from step 3. By correlating each input data, an analysis result is obtained that determines whether an abnormal situation exists and its level of urgency.
[0748] Step 5:
[0749] The server generates notifications via communication functions based on the analysis results of abnormal situations and urgency levels from the intelligent module. These notifications include the user's emotional state and environmental information, and are sent to the relevant administrator terminal as emergency notifications when a rapid response is required.
[0750] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0751] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0752] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0753] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0754] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0755] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0756] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0757] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0758] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0759] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0760] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0761] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0762] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0763] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0764] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0765] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0766] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0767] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0768] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0769] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0770] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0771] The following is further disclosed regarding the embodiments described above.
[0772] (Claim 1)
[0773] A means for a voice recognition sensor to detect a predetermined keyword,
[0774] A sensor means for collecting environmental data to detect abnormal values,
[0775] An artificial intelligence module for analyzing collected data,
[0776] A communication means for generating and sending notifications based on analysis results,
[0777] A system that includes this.
[0778] (Claim 2)
[0779] The system according to claim 1, further comprising means for sending an emergency notification to a registered terminal when the analysis results indicate an abnormality that exceeds a predetermined standard.
[0780] (Claim 3)
[0781] The system according to claim 1, further comprising a function to receive user feedback and adjust system parameters.
[0782] "Example 1"
[0783] (Claim 1)
[0784] A voice recognition means for detecting predetermined keywords,
[0785] A means of collecting environmental information to monitor the surrounding conditions and detect anomalies,
[0786] An analysis means equipped with an intelligent module for processing the generated data,
[0787] A communication means that creates and sends a warning based on the analysis results,
[0788] A means of receiving information from users and adjusting system settings,
[0789] A system that includes this.
[0790] (Claim 2)
[0791] The system according to claim 1, further comprising means for sending an emergency warning to a registered information terminal when the analysis result indicates an abnormality that exceeds a predetermined threshold.
[0792] (Claim 3)
[0793] The system according to claim 1, comprising a function to receive user feedback and improve the analysis accuracy of the intelligent module.
[0794] "Application Example 1"
[0795] (Claim 1)
[0796] A means for a speech recognition detection device to detect specific instructional words,
[0797] A sensor device for measuring external environmental values,
[0798] A machine learning module means for evaluating acquired information,
[0799] A communication device for creating and distributing notices based on the analyzed results,
[0800] A notification function means for issuing alerts in response to emergencies,
[0801] A system that includes this.
[0802] (Claim 2)
[0803] The system according to claim 1, which includes means for determining an abnormal situation within a public facility based on the analysis results and sending an emergency notification to registered user terminals.
[0804] (Claim 3)
[0805] The system according to claim 1, comprising a function to receive responses from users and optimize evaluation metrics.
[0806] "Example 2 of combining an emotion engine"
[0807] (Claim 1)
[0808] A means by which a speech recognition device monitors a specified keyword,
[0809] A collection device for collecting environmental information and detecting anomalies,
[0810] A machine learning module for detailed evaluation of the analyzed data,
[0811] A communication device for generating and transmitting adaptive notifications based on evaluation results,
[0812] An emotion recognition engine for identifying and analyzing emotional states,
[0813] A system that includes this.
[0814] (Claim 2)
[0815] The system according to claim 1, comprising means for sending an emergency notification to registered devices when an abnormality exceeding a specified standard is detected.
[0816] (Claim 3)
[0817] The system according to claim 1, comprising a function to receive feedback from users and optimize system operation.
[0818] "Application example 2 when combining with an emotional engine"
[0819] (Claim 1)
[0820] A means for a speech recognition device to detect a predetermined emergency term,
[0821] A means of analyzing a user's emotional state using an emotion recognition engine,
[0822] A device for collecting environmental information to detect abnormal values,
[0823] Intelligent module functions for analyzing collected information,
[0824] A communication function for generating and sending notifications based on analysis results and emotional state,
[0825] A system that includes this.
[0826] (Claim 2)
[0827] The system according to claim 1, which includes a function to send an emergency notification to a registered terminal when the analysis results or emotional state indicate an abnormality that exceeds a predetermined standard.
[0828] (Claim 3)
[0829] The system according to claim 1, comprising a function to receive user feedback and adjust system settings. [Explanation of Symbols]
[0830] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for a voice recognition sensor to detect a predetermined keyword, A sensor means for collecting environmental data to detect abnormal values, An artificial intelligence module for analyzing collected data, A communication means for generating and sending notifications based on analysis results, A system that includes this.
2. The system according to claim 1, further comprising means for sending an emergency notification to a registered terminal when the analysis results indicate an abnormality that exceeds a predetermined standard.
3. The system according to claim 1, further comprising a function to receive user feedback and adjust system parameters.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A