system

A system using facial and voice analysis with an emotion engine provides real-time risk assessment and tailored responses to enhance residential security by accurately evaluating visitor risk and addressing user emotions.

JP2026103524APending Publication Date: 2026-06-24SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-12
Publication Date
2026-06-24

Smart Images

  • Figure 2026103524000001_ABST
    Figure 2026103524000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A receiving means that acquires facial images and audio in order to identify the person visiting, An evaluation method that analyzes acquired facial images and audio to assess the risk level of visitors, A response system that automatically responds to visitors based on the assessed risk level, A communication method for notifying users of visitor information and risk assessment, A remote monitoring system that uses facial images and audio to assess the safety level of visitors in real time, and allows users to check the level of risk via their mobile devices. A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, suspicious behavior by visitors has been increasing, and ensuring the safety of residents, especially the elderly living alone, has become a social issue. However, existing security systems have problems in accurately evaluating the risk level of visitors and taking appropriate actions, and there is insufficient early warning and response to suspicious persons. In addition to these, there is a lack of means for direct response when residents are in a remote location, and technologies for enhancing safety are required.

Means for Solving the Problems

[0005] This invention acquires facial images and voice recordings of visitors, compares them with a database to identify visitors and assess their risk level. Based on the assessed risk level, the system automatically responds to visitors and notifies the user in real time. This allows for remote response even when residents are not present, enabling early detection and rapid response to suspicious individuals. Furthermore, the system allows users to remotely notify the police based on visitor information and assessments, effectively enhancing resident safety.

[0006] "Acquisition means" refers to a device or function for collecting a visitor's facial image and voice.

[0007] "Evaluation method" refers to a function that analyzes acquired facial images and audio and compares them with a database to calculate the visitor's risk level.

[0008] "Response mechanism" refers to a function that automatically presents appropriate messages or signals to visitors based on their assessed risk level.

[0009] "Notification means" refers to a function that transmits visitor information and risk assessment results to the user in real time, enabling the user to remotely issue response instructions.

[0010] A "database" refers to a collection of information, such as visitors' facial images, past visit history, and criminal lists, that is stored and used for reference by evaluation tools.

[0011] "Risk level" refers to an assessment result that expresses potential risks numerically or categorized, based on information about the visitor. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, a communication I / F (Interface) with a reference numeral is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] This invention provides a system that acquires and analyzes visitor information and assesses risk levels to provide appropriate responses to visitors. This system is operated by integrating terminals, servers, and users.

[0034] The terminal is installed at the building's entrance and is equipped with a camera and microphone to capture images of visitors' faces and record their voices. The terminal transmits this data to a server in real time.

[0035] The server uses facial recognition and voice analysis technologies to analyze data transmitted from the terminal. The server compares the visitor's face to known facial data in the database and checks for matches with visit history and criminal lists. Furthermore, the server infers emotions and intentions from the visitor's voice and scores the visitor's risk level.

[0036] Based on the risk level assessed by the server, the terminal automatically responds to visitors. Responses can be pre-set messages or voice messages via the intercom. The terminal can also sound an alarm if necessary.

[0037] Users can access the system in real time from their smartphones or computers via the internet. They can view visitor information and risk assessments provided by the server, and, if necessary, respond remotely or instruct the police to be notified. This allows users to respond appropriately to visitors even when they are away, ensuring the safety of their homes.

[0038] For example, when a visitor appears at the front door, the device automatically takes a picture of their face and records their voice. The server analyzes this data, and if it determines that the visitor is a legitimate delivery person with a history of no problems, the device plays an automated message such as, "Thank you for the delivery, please leave your package at the front door." On the other hand, if the server detects any suspicious behavior from the visitor, it issues an alarm to warn the user, and a notification is immediately sent to the user. After reviewing the situation, the user can, if necessary, report it to the police or share the details of the situation with their family.

[0039] Thus, the system of the present invention enables a quick and accurate response to visitors and serves as an effective means of enhancing the security function of a residence.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0043] Step 2:

[0044] The device transmits facial image data and audio data acquired by the device to the server in real time. The server receives this data.

[0045] Step 3:

[0046] The server analyzes facial images using a facial recognition algorithm and compares them with records in the database. In particular, it checks for matches with past visit history and criminal lists.

[0047] Step 4:

[0048] The server analyzes the audio data. It analyzes the characteristics of the voice and applies speech processing techniques to estimate the speaker's emotions and intentions.

[0049] Step 5:

[0050] The server integrates facial recognition results and voice analysis results to score the visitor's risk level. This provides a safety assessment of the visitor.

[0051] Step 6:

[0052] The server sends an appropriate automated response instruction to the terminal based on the risk assessment. The terminal receives this instruction and sends a response message to the visitor.

[0053] Step 7:

[0054] If the device determines that a visitor poses a high risk, it will sound an alarm to alert those nearby. Additionally, a danger notification will be immediately sent from the server to the user.

[0055] Step 8:

[0056] Users receive notifications, access them via the internet, and verify visitor information and risk assessments. Remote responses or instructions to contact the police are provided as needed.

[0057] Step 9:

[0058] The server receives instructions from the user and automatically notifies the police. It also updates the database and stores information to be used for future visitor detection.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] In recent years, the need for home security has been increasing. However, there is a problem in quickly and accurately understanding the intentions of visitors and responding appropriately, whether the homeowner is at home or away. Furthermore, existing security systems lack the ability to quantitatively assess the risk posed by visitors and provide appropriate real-time notifications to users. This invention aims to solve these problems.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for acquiring images and sounds using an imaging device to identify a visitor, means for evaluating the visitor's level of caution by processing the acquired images and sounds, and means for transmitting visitor information and caution level information to the user. This makes it possible to quickly analyze the visitor's intentions and level of caution and to immediately provide the user with necessary notifications.

[0064] An "imaging device" is a device used to acquire data such as images and sounds, and generally includes cameras and microphones.

[0065] An "image" is data that represents visual information such as a visitor's face or body.

[0066] "Acoustics" refers to data that represents auditory information, including visitors' voices and ambient sounds.

[0067] "Processing" refers to a series of actions taken to analyze acquired image and audio data and derive specific information.

[0068] "Alert level" is an indicator that quantifies or assigns numerical values ​​to evaluate the level of risk a visitor poses.

[0069] An "information recording device" is a device that stores data and lists of known visitors and verifies that information as needed.

[0070] A "user" is an individual or organization that operates this system and interacts with visitors.

[0071] "Notifications" refer to actions or functions that inform users about visitor information and the level of alertness.

[0072] "Sequential delivery" means providing information in real time by sending information continuously without any time gaps.

[0073] "Response instructions" refer to instructions or reactions that users give to visitors through the system.

[0074] This invention is a system that acquires a visitor's facial image and voice, analyzes this data to evaluate the visitor's level of caution, and provides an appropriate response to the visitor immediately. This system consists of a terminal, a server, and a user.

[0075] The terminal is installed at the building's entrance and used to capture images and audio of visitors. A high-resolution imaging device (camera) and a high-sensitivity audio input device (microphone) are used to quickly capture the visitor's facial and voice characteristics and transmit them as digital data to a server.

[0076] The server is the central device for processing the acquired data. Image data processing uses image analysis libraries such as OpenCV to identify visitors' faces by comparing them to data in the database. Furthermore, audio processing modules such as Google Cloud Speech-to-Text API are used for speech analysis to estimate emotions and intentions from the visitor's voice. Based on these analysis results, the server scores the visitor's level of caution, enabling real-time decision-making.

[0077] Based on data processing by the server, if a visitor is deemed high-risk, the terminal will either play an automated response message or issue an alarm. This ensures a safe response without direct interaction.

[0078] Users can access the system via the internet to check visitor information and alert levels in real time. Using smart devices, users can receive notifications and respond appropriately no matter where they are. This allows users to maintain the security of their homes with peace of mind, even when they are away.

[0079] As a concrete example, when a visitor appears at the front door, the terminal takes a picture of the person with its camera and records their voice. The server analyzes this data, and if it recognizes, for example, the person as a delivery person who has visited the house before, the terminal will automatically respond with, "Thank you for the delivery, please leave the package at the front door." Furthermore, if there is anything suspicious about the visitor's behavior, the terminal will sound an alarm and immediately notify the user. An example of a prompt to the generating AI model in this case would be, "Please explain how the system works to analyze the visitor's facial image and voice data to assess the level of caution. Also, please give specific examples of responses."

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The device captures a facial image with its camera and records audio with its microphone when a visitor arrives at the building entrance. The input consists of visual and audio information of the visitor's face, and the output is the storage of this data in a digital format. This process utilizes a high-resolution camera and a high-sensitivity microphone, and is designed to adapt to changes in the external environment.

[0083] Step 2:

[0084] The device transmits acquired facial images and audio data to the server in real time. The input is the digital data collected in step 1, and the output is data transmitted to the server in an encrypted format. This process uses a secure communication protocol to protect the confidentiality of the data.

[0085] Step 3:

[0086] The server processes the received facial image data through a facial recognition algorithm. The input data is the transmitted facial image, and the output is the result of matching it with a known facial database. Specifically, the OpenCV library is used here to determine whether the visitor is a registered person.

[0087] Step 4:

[0088] The server processes audio data using a speech analysis module. The input is the transmitted audio data, and the output is the analysis results of emotions and intentions extracted from the audio. The Google Cloud Speech-to-Text API is used to convert the audio to text, and then sentiment analysis is performed to evaluate the visitor's psychological state.

[0089] Step 5:

[0090] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. The input is the results of facial recognition and voice analysis, and the output is the visitor's risk level score. This score is determined based on criteria defined by the server and assesses the overall risk level of the visitor.

[0091] Step 6:

[0092] The terminal plays an automated response message or issues an alarm based on the alert level score received from the server. The input is the alert level score calculated in step 5, and the output is a voice message or alarm for the visitor. Specifically, if the visitor is friendly, a message such as "Welcome" is played, and if there is a reason to be cautious, an alarm sounds immediately.

[0093] (Application Example 1)

[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] In modern society, ensuring the safety of visitors to homes and facilities is a critical issue. Furthermore, there is a need to quickly analyze visitor information and take appropriate action. However, existing security systems struggle to assess visitor risk in real time, making it difficult for users to effectively manage them remotely. Therefore, there is a need for the development of a system that accurately analyzes visitor risk and allows users to take immediate action.

[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0097] In this invention, the server includes receiving means for identifying visitors, evaluation means for analyzing acquired data to assess the visitor's risk level, and communication means for notifying the user of the visitor's information and risk level assessment. This enables real-time assessment of the visitor's safety level, allowing the user to check the risk level via a mobile device and respond quickly.

[0098] "Receiving means" refers to a device for acquiring facial images and voices of visitors.

[0099] "Evaluation means" refers to a device or method for analyzing facial images and audio acquired by a receiving means to evaluate the level of risk a visitor poses.

[0100] "Communication means" refers to a device or method for notifying users of visitor information and risk assessments.

[0101] "Remote monitoring means" refers to a device or method that enables a user to check the safety level of a visitor in real time via a portable electronic device.

[0102] This invention relates to a security system that analyzes visitor information in real time to enhance the safety of a residence. A server receives facial images and voice data from a terminal installed at the building's entrance to identify visitors. Facial images are analyzed using the "face_recognition" library and compared against known facial data in a database. Voice data is converted to text using the "speech_recognition" library, and its emotions are analyzed.

[0103] Based on this data, the server assesses the visitor's risk level and notifies the terminal or the user's portable electronic device of the risk level via communication means. Users can check the visitor's safety level in real time via their smartphone or other mobile device and take prompt and appropriate action as needed.

[0104] For example, when a delivery person arrives, the server recognizes their face, and if it determines from past history that they are safe, it plays a message on the device such as, "Thank you for the delivery, please leave your package at the doorstep." In the unlikely event that the visitor is deemed suspicious, an alarm is issued and the user is immediately notified.

[0105] An example of a prompt message would be, "When a visitor arrives, we will capture their photo and audio to determine if they are known or suspicious. If they are safe, we will send a notification to your smartphone." In this way, users can effectively manage the security of their homes and facilities.

[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0107] Step 1:

[0108] The device captures the visitor's facial image and voice.

[0109] The device uses a camera and microphone to capture facial images and audio data of visitors. This data is transmitted to the server in real time. The input is raw image and audio data, and the output is the transmission of data to the server.

[0110] Step 2:

[0111] The server analyzes the facial image data.

[0112] The server analyzes the received face image data using the "face_recognition" library. It compares it with known face data in the database to determine if there is a match. The input is face image data, and the output is the match result and identification information.

[0113] Step 3:

[0114] The server analyzes the audio data.

[0115] The server uses the "speech_recognition" library to convert speech data into text. Based on this text data, it performs sentiment analysis to infer the visitor's intent. The input is speech data, and the output is text data and the sentiment analysis results.

[0116] Step 4:

[0117] The server assesses the level of risk.

[0118] The server scores the visitor's risk level based on facial recognition and sentiment analysis results. This evaluation takes into account the visitor's past history and emotional state. The input is the facial recognition result and sentiment analysis result, and the output is the risk level score.

[0119] Step 5:

[0120] Notify the device and the user.

[0121] Based on the assessed risk score, the server either plays an audio message to the visitor via the terminal or sends a notification to the user's mobile device. The input is the risk score, and the output is the automated response on the terminal or the notification to the user.

[0122] Step 6:

[0123] The user selects the next action.

[0124] Users check the information conveyed via their smartphones and take action as needed, such as reporting to the police or notifying family members. The input is the notification information from the system, and the output is the specific action to be taken.

[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0126] This invention combines a system that uses facial images and voice to assess visitor safety with an emotion engine that recognizes user emotions. This system functions in conjunction with the terminal, server, user, and emotion engine.

[0127] The device is installed at the entrance of the house and is equipped with a camera to capture images of visitors' faces and a microphone to record their voices. The acquired facial images and voice data are transmitted to a server in real time.

[0128] The server matches received facial images against a database to identify visitors, estimates their emotions through voice analysis, and assesses their risk level. The emotion engine also analyzes user voice and touch actions to recognize their emotions. This enables the presentation of information and responses tailored to the user's emotional state.

[0129] For example, if the emotion engine detects anxiety from the user's voice, the server can simplify the display of visitor information and provide additional reassuring information to alleviate the user's anxiety. Furthermore, it can recommend automated response options to ensure the user can respond comfortably.

[0130] As a concrete example, when a visitor arrives, the device retrieves the visitor's information, and the server assesses the level of risk. Once the assessment results are sent, the emotion engine recognizes the user's emotions. If the user is feeling anxious, the emotion engine instructs the system to display reassuring information to the user, presenting them with the message, "This visitor is safe." It can also recommend example responses that will help the user interact with the visitor in a friendly manner. This allows the user to interact with visitors with confidence.

[0131] By incorporating an emotion engine in this way, it becomes possible to go beyond simply assessing the risk level of visitors and instead provide detailed responses based on the user's emotions, thereby improving their sense of security.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0135] Step 2:

[0136] The device sends the acquired facial image data and audio data to the server. The server receives this data in real time.

[0137] Step 3:

[0138] The server analyzes the received facial images using a facial recognition algorithm, compares them with records in the database, and identifies the visitor. In particular, it checks for matches with past visit history and criminal lists.

[0139] Step 4:

[0140] The server analyzes the audio data using speech analysis technology. It estimates the speaker's emotions and intentions from the audio and analyzes the visitor's level of caution and trustworthiness.

[0141] Step 5:

[0142] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. This allows for an assessment of the visitor's safety.

[0143] Step 6:

[0144] The server sends instructions to the terminal based on the risk assessment. The terminal then executes an appropriate automated response to the visitor based on those instructions.

[0145] Step 7:

[0146] If a device is identified as a high-risk visitor, an alarm sounds to warn those nearby of the danger, and a danger notification is immediately sent from the server to the user.

[0147] Step 8:

[0148] The emotion engine analyzes the user's voice and touch actions to recognize the user's emotional state. Based on the user's emotions, it optimizes information presentation and response options.

[0149] Step 9:

[0150] Users can access the system via smartphones or computers to view visitor information and ratings. They can then use responses suggested by the emotion engine to remotely address situations or instruct police to be notified.

[0151] Step 10:

[0152] The server receives instructions from users and takes action, such as automatically reporting issues. It also adds evaluation results and visitor information to a database to record data for future reference.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] Conventional visitor identification systems only assess the safety of visitors and lack consideration for the user's emotional state. Specifically, they did not provide guidance on how users should respond to visitors, making it difficult to respond appropriately. Furthermore, accurately assessing the visitor's emotions was difficult when determining their level of caution.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes a device for acquiring images and audio to identify visitors, a device for analyzing the acquired images and audio to determine the visitor's level of alertness, and a device for detecting the user's emotional state and adjusting the information presented. This enables not only the assessment of the visitor's safety but also flexible responses and the provision of safety information in accordance with the user's emotional state.

[0158] "Image" refers to visual information used to visually record and analyze the faces and appearances of visitors.

[0159] "Audio" refers to auditory information used to record and analyze sounds emitted by visitors and users.

[0160] "Device" refers to a combination of hardware and software used to acquire, analyze, and display data.

[0161] "Alert level" refers to a measure of safety that is assessed based on the behavior and emotions of visitors.

[0162] "User emotional state" refers to the results of analyzing the emotional responses a user exhibits from their voice and behavior.

[0163] "Information presentation" refers to guidance and messages displayed to the user based on the analysis results.

[0164] This system identifies visitors, assesses their safety, and provides appropriate responses based on the user's emotional state. The system operates through the collaborative efforts of a terminal, a server, and an engine with emotional assessment capabilities.

[0165] Device features:

[0166] The device is installed at the entrance of the house and, when a visitor appears, captures their facial image with a camera and records their voice with a microphone. This data is transmitted to a server via secure communication. The device can be implemented, for example, using a camera module based on a Raspberry Pi and a high-sensitivity microphone for voice acquisition.

[0167] Server functions:

[0168] The server performs facial recognition by comparing received facial images with internal data. It can utilize facial recognition algorithms such as FaceNet or the OpenCV library. For audio data, it performs speech analysis using the Google Cloud Speech-to-Text API to estimate the visitor's emotions. This allows for a comprehensive evaluation of the visitor's level of caution. Based on this evaluation, information is provided to the user in real time.

[0169] Emotional engine function:

[0170] The user's emotional state is evaluated by an emotion engine based on voice tone and device touch operations. Real-time deep learning is utilized, such as with NVIDIA Jetson, to achieve emotion analysis. Based on the emotion engine's analysis results, information and responses tailored to the user's state are integrated.

[0171] Specific example:

[0172] When a visitor arrives at the front door, the device sends a facial image and audio to the server. The server determines the visitor is safe, and even if the emotion engine detects anxiety from the user's voice, it provides the user with reassuring information.

[0173] Example of a prompt:

[0174] "The visitor's facial image and audio data have been sent to the server. Please evaluate whether this visitor is safe and, based on that information, display a reassuring message to the user."

[0175] The above system enables visitor safety assessment and user support, providing users with a sense of security.

[0176] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0177] Step 1:

[0178] When the terminal detects a visitor's presence, it captures a facial image with its camera and records audio with its microphone. This data generates facial information as an image and audio information, which are then sent to the server in digital format. Specifically, a camera module using a Raspberry Pi captures an image, and an audio device captures clear audio.

[0179] Step 2:

[0180] The server processes the received facial image through a facial recognition algorithm and compares it with existing facial data within the system. The input is the transmitted facial image, and the output is the result of the facial recognition. This determines whether the visitor is a person who has been identified in the past. This process is performed using facial recognition technology such as FaceNet.

[0181] Step 3:

[0182] Simultaneously, the server inputs the audio data into a speech sentiment analysis tool to perform calculations that estimate the visitor's emotions. The input is the audio data, and the output is the estimated emotional state. The Google Cloud Speech-to-Text API and sentiment analysis libraries are used for speech-to-text conversion and emotion estimation.

[0183] Step 4:

[0184] The server comprehensively evaluates the visitor's level of caution based on the results of facial recognition and voice emotion estimation. The input is the facial recognition result and the voice emotion evaluation result, and the output is the caution level. This evaluation is a process that quantifies how much of a risk the visitor poses to the user.

[0185] Step 5:

[0186] The emotion engine analyzes the user's emotions. It analyzes the user's emotional state based on their voice and device interactions. Input is the user's voice or touch data, and output is the detected emotional state. Data processing using NVIDIA Jetson enables real-time estimation of the user's emotions.

[0187] Step 6:

[0188] The server generates information to guide the user based on the visitor's alert level and the user's emotional state. The input is the visitor's alert level and the user's emotional state, and the output is a message directed at the user. A generative AI model is used to provide specific messages and response examples as needed.

[0189] Each step works in conjunction with the others to create a system that supports visitor safety assessment and user interaction.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0192] In recent years, with the increase in visitors, there has been a growing need for an efficient system that can quickly and accurately assess visitor safety and provide users with a sense of security. Conventional systems are limited to providing information about visitors and do not provide appropriate information tailored to the user's emotional state, leaving the possibility of users experiencing tension or anxiety. This invention aims to solve such problems and improve the sense of security users experience when interacting with visitors.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes means for acquiring images and sounds to identify visitors, means for analyzing the acquired data to evaluate the visitor's safety, and means for adjusting the displayed content based on the user's emotional state and providing reassuring information. This makes it possible to provide users with immediate and appropriate information regarding the visitor's safety, thereby improving their sense of security.

[0195] A "visitor" is an individual whose facial image and voice are acquired by the system, and whose safety and emotional state are evaluated.

[0196] "Images" are visual data acquired to visually capture the characteristics of visitors.

[0197] "Acoustics" refers to auditory data acquired to capture visitors' voices and other auditory characteristics.

[0198] "Means" refer to the structure or process established by a system to achieve a specific function.

[0199] A "server" is a computer device that receives and analyzes acquired data and processes information in cooperation with other devices and systems.

[0200] A "user" is someone who receives visitor information through the system and provides instructions or responses to visitors.

[0201] "Safety" is the result of an assessment that visitors do not pose a potential risk based on the data collected.

[0202] The system of the present invention includes a sophisticated process for evaluating the safety of visitors and providing users with a sense of security. The terminal collects the visitor's facial image and voice and transmits them to the server. The server analyzes the received data and performs a safety evaluation. Specifically, it matches the facial image with a storage medium and analyzes the voice to identify emotions.

[0203] The server then uses a generative AI model to analyze the user's emotional state and, if necessary, generates information to provide reassurance. For example, it can display a message such as "This visitor is safe" to a user who is feeling uneasy about the visitor.

[0204] The software used includes machine learning libraries such as TENSORFLOW® for facial image recognition and the Google Cloud Speech-to-Text API for speech data analysis. IBM Watson® Tone Analyzer is also used to assess user emotions. This allows the system to instantly present information appropriate to the user's state.

[0205] As a concrete example, when a delivery person visits a home, the terminal captures the delivery person's face and voice. The server analyzes this data to confirm that the delivery person's face is registered in the database and detects a friendly tone of voice. If the user is feeling stressed, the system displays a message such as "The delivery person is safe and reliable" to enhance the user's sense of security.

[0206] An example of a prompt message might be, "Assess the visitor's safety and emotional state based on their facial image and voice, and display information to alleviate their anxiety." This process allows users to interact with visitors with confidence.

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The device acquires facial images and audio from visitors. It uses raw data obtained from the camera and microphone as input. This data is recorded as image and audio data.

[0210] Step 2:

[0211] The terminal sends the acquired facial image and audio data to the server. The input is the data acquired in step 1, and the output is digital data converted into a format that can be received by the server.

[0212] Step 3:

[0213] The server identifies the visitor by comparing the received facial image with the storage medium. Using the facial image data as input, it performs a database search and outputs the visitor's ID and related information as a result.

[0214] Step 4:

[0215] The server converts audio data into text using the Google Cloud Speech-to-Text API and performs sentiment analysis based on that text data. The input is audio data, and the output is an evaluation result indicating the emotional state.

[0216] Step 5:

[0217] The server uses IBM Watson Tone Analyzer to analyze the user's emotional state. The input is a visitor safety assessment generated internally by the server, and the output determines the reassuring information that should be provided to the user.

[0218] Step 6:

[0219] The server sends data to the user's terminal to display visitor information and safety assessment results. The input is the results of steps 3 and 5, and the output is the message to be displayed and the recommended action.

[0220] Step 7:

[0221] The terminal displays information received from the server to the user and informs them of the system's final result. Input is data from the server, and output is information displayed on the terminal's screen. This allows the user to respond appropriately to visitors.

[0222] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0223] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0224] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0225] [Second Embodiment]

[0226] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0227] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0228] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0229] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0230] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0232] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0233] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0234] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0235] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0236] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0237] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0238] This invention provides a system that acquires and analyzes visitor information and assesses risk levels to provide appropriate responses to visitors. This system is operated by integrating terminals, servers, and users.

[0239] The terminal is installed at the building's entrance and is equipped with a camera and microphone to capture images of visitors' faces and record their voices. The terminal transmits this data to a server in real time.

[0240] The server uses facial recognition and voice analysis technologies to analyze data transmitted from the terminal. The server compares the visitor's face to known facial data in the database and checks for matches with visit history and criminal lists. Furthermore, the server infers emotions and intentions from the visitor's voice and scores the visitor's risk level.

[0241] Based on the risk level assessed by the server, the terminal automatically responds to visitors. Responses can be pre-set messages or voice messages via the intercom. The terminal can also sound an alarm if necessary.

[0242] Users can access the system in real time from their smartphones or computers via the internet. They can view visitor information and risk assessments provided by the server, and, if necessary, respond remotely or instruct the police to be notified. This allows users to respond appropriately to visitors even when they are away, ensuring the safety of their homes.

[0243] For example, when a visitor appears at the front door, the device automatically takes a picture of their face and records their voice. The server analyzes this data, and if it determines that the visitor is a legitimate delivery person with a history of no problems, the device plays an automated message such as, "Thank you for the delivery, please leave your package at the front door." On the other hand, if the server detects any suspicious behavior from the visitor, it issues an alarm to warn the user, and a notification is immediately sent to the user. After reviewing the situation, the user can, if necessary, report it to the police or share the details of the situation with their family.

[0244] Thus, the system of the present invention enables a quick and accurate response to visitors and serves as an effective means of enhancing the security function of a residence.

[0245] The following describes the processing flow.

[0246] Step 1:

[0247] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0248] Step 2:

[0249] The device transmits facial image data and audio data acquired by the device to the server in real time. The server receives this data.

[0250] Step 3:

[0251] The server analyzes facial images using a facial recognition algorithm and compares them with records in the database. In particular, it checks for matches with past visit history and criminal lists.

[0252] Step 4:

[0253] The server analyzes the audio data. It analyzes the characteristics of the voice and applies speech processing techniques to estimate the speaker's emotions and intentions.

[0254] Step 5:

[0255] The server integrates facial recognition results and voice analysis results to score the visitor's risk level. This provides a safety assessment of the visitor.

[0256] Step 6:

[0257] The server sends an appropriate automated response instruction to the terminal based on the risk assessment. The terminal receives this instruction and sends a response message to the visitor.

[0258] Step 7:

[0259] If the device determines that a visitor poses a high risk, it will sound an alarm to warn those nearby. Additionally, a danger notification will be immediately sent from the server to the user.

[0260] Step 8:

[0261] The user receives a notification, accesses it via the internet, and verifies visitor information and risk assessment. If necessary, remote responses or instructions to report to the police are provided.

[0262] Step 9:

[0263] The server receives instructions from the user and automatically notifies the police. It also updates the database and stores information to be used for future visitor detection.

[0264] (Example 1)

[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] In recent years, the need for home security has been increasing. However, there is a problem in quickly and accurately understanding the intentions of visitors and responding appropriately, whether the homeowner is at home or away. Furthermore, existing security systems lack the ability to quantitatively assess the risk posed by visitors and provide appropriate real-time notifications to users. This invention aims to solve these problems.

[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0268] In this invention, the server includes means for acquiring images and sounds using an imaging device to identify a visitor, means for evaluating the visitor's level of caution by processing the acquired images and sounds, and means for transmitting visitor information and caution level information to the user. This makes it possible to quickly analyze the visitor's intentions and level of caution and to immediately provide the user with necessary notifications.

[0269] An "imaging device" is a device used to acquire data such as images and sounds, and generally includes cameras and microphones.

[0270] An "image" is data that represents visual information such as a visitor's face or body.

[0271] "Acoustics" refers to data that represents auditory information, including visitors' voices and ambient sounds.

[0272] "Processing" refers to a series of actions taken to analyze acquired image and audio data and derive specific information.

[0273] "Alert level" is an indicator that quantifies or assigns numerical values ​​to evaluate the level of risk a visitor poses.

[0274] An "information recording device" is a device that stores data and lists of known visitors and verifies that information as needed.

[0275] A "user" is an individual or organization that operates this system and interacts with visitors.

[0276] "Notifications" refer to actions or functions that inform users about visitor information and the level of alertness.

[0277] "Sequential delivery" means providing information in real time by sending information continuously without any time gaps.

[0278] "Response instructions" refer to instructions or reactions that users give to visitors through the system.

[0279] This invention is a system that acquires a visitor's facial image and voice, analyzes this data to evaluate the visitor's level of caution, and provides an appropriate response to the visitor immediately. This system consists of a terminal, a server, and a user.

[0280] The terminal is installed at the building's entrance and used to capture images and audio of visitors. A high-resolution imaging device (camera) and a high-sensitivity audio input device (microphone) are used to quickly capture the visitor's facial and voice characteristics and transmit them as digital data to a server.

[0281] The server is the central device for processing the acquired data. Image data processing uses image analysis libraries such as OpenCV to identify visitors' faces by comparing them to data in the database. Furthermore, audio processing modules such as the Google Cloud Speech-to-Text API are used for speech analysis to estimate emotions and intentions from the visitor's voice. Based on these analysis results, the server scores the visitor's level of caution, enabling real-time decision-making.

[0282] Based on data processing by the server, if a visitor is deemed high-risk, the terminal will either play an automated response message or issue an alarm. This ensures a safe response without direct interaction.

[0283] Users can access the system via the internet to check visitor information and alert levels in real time. Using smart devices, users can receive notifications and respond appropriately no matter where they are. This allows users to maintain the security of their homes with peace of mind, even when they are away.

[0284] As a specific example, when a visitor appears in front of the entrance, the terminal takes a picture of the person with a camera and records the voice. The server analyzes this. For example, if it is recognized that the person is a delivery person who has visited the house in the past, the terminal automatically responds with "Thank you for the delivery. Please leave the package at the entrance." Also, if there is anything suspicious about the behavior of the visitor, the terminal sounds an alarm and immediately notifies the user. As an example of the prompt text for the generative AI model at this time, "Please explain how a system that analyzes the face image and voice data of a visitor and evaluates the degree of alertness functions. Also, give specific examples of responses." is used.

[0285] The flow of the specific process in Example 1 will be described using FIG. 11.

[0286] Step 1:

[0287] When the visitor arrives at the entrance of the building, the terminal takes a face image with a camera and records the voice with a microphone. The inputs are the visual information and voice information of the visitor's face, and the output is that these data are saved in digital format. In this process, a high-resolution camera and a high-sensitivity microphone are used, and it is designed to be adaptable to changes in the external environment.

[0288] Step 2:

[0289] The terminal sends the acquired face image and voice data to the server in real time. The input is the digital data collected in Step 1, and the output is the data transmitted to the server in an encrypted format. In this process, a secure communication protocol is used to protect the confidentiality of the data.

[0290] Step 3:

[0291] The server processes the received facial image data through a facial recognition algorithm. The input data is the transmitted facial image, and the output is the result of matching it with a known facial database. Specifically, the OpenCV library is used here to determine whether the visitor is a registered person.

[0292] Step 4:

[0293] The server processes audio data using a speech analysis module. The input is the transmitted audio data, and the output is the analysis results of emotions and intentions extracted from the audio. The Google Cloud Speech-to-Text API is used to convert the audio to text, and then sentiment analysis is performed to evaluate the visitor's psychological state.

[0294] Step 5:

[0295] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. The input is the results of facial recognition and voice analysis, and the output is the visitor's risk level score. This score is determined based on criteria defined by the server and assesses the overall risk level of the visitor.

[0296] Step 6:

[0297] The terminal plays an automated response message or issues an alarm based on the alert level score received from the server. The input is the alert level score calculated in step 5, and the output is a voice message or alarm for the visitor. Specifically, if the visitor is friendly, a message such as "Welcome" is played, and if there is a reason to be cautious, an alarm sounds immediately.

[0298] (Application Example 1)

[0299] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0300] In modern society, ensuring the safety of visitors to homes and facilities is a critical issue. Furthermore, there is a need to quickly analyze visitor information and take appropriate action. However, existing security systems struggle to assess visitor risk in real time, making it difficult for users to effectively manage them remotely. Therefore, there is a need for the development of a system that accurately analyzes visitor risk and allows users to take immediate action.

[0301] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0302] In this invention, the server includes receiving means for identifying visitors, evaluation means for analyzing acquired data to assess the visitor's risk level, and communication means for notifying the user of the visitor's information and risk level assessment. This enables real-time assessment of the visitor's safety level, allowing the user to check the risk level via a mobile device and respond quickly.

[0303] "Receiving means" refers to a device for acquiring facial images and voices of visitors.

[0304] "Evaluation means" refers to a device or method for analyzing facial images and audio acquired by a receiving means to evaluate the level of risk a visitor poses.

[0305] "Communication means" refers to a device or method for notifying users of visitor information and risk assessments.

[0306] "Remote monitoring means" refers to a device or method that enables a user to check the safety level of a visitor in real time via a portable electronic device.

[0307] This invention relates to a security system that analyzes visitor information in real time to enhance the security of a house. The server receives face images and voice data from a terminal installed at the entrance of a building in order to identify the visitor. The face image is analyzed using the "face_recognition" library and compared with known face data in the database. The voice data is converted into text using the "speech_recognition" library, and its sentiment is analyzed.

[0308] Based on this data, the server evaluates the risk level of the visitor and communicates the risk level information to the terminal or the user's portable electronic device through communication means. The user can confirm the safety level of the visitor in real time through a smartphone or other portable device and can take prompt and appropriate actions if necessary.

[0309] As a specific example, when a delivery person visits, if the server recognizes the face and determines that it is safe based on past history, it plays a message such as "Thank you for the delivery. Please leave the package at the entrance" on the terminal. Also, if the visitor is regarded as suspicious, an alarm is issued and the user is immediately notified.

[0310] Examples of prompt sentences include "When a visitor comes, acquire the face photo and voice, and determine whether it is known or suspicious. If it is safe, send a notification to the smartphone." In this way, the user can effectively manage the security of their home or facility.

[0311] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0312] Step 1:

[0313] ​​​​​The device uses a camera and microphone to capture facial images and audio data of visitors. This data is transmitted to the server in real time. The input is raw image and audio data, and the output is the transmission of data to the server.

[0315] Step 2:

[0316] The server analyzes the facial image data.

[0317] The server analyzes the received face image data using the "face_recognition" library. It compares it with known face data in the database to determine if there is a match. The input is face image data, and the output is the match result and identification information.

[0318] Step 3:

[0319] The server analyzes the audio data.

[0320] The server uses the "speech_recognition" library to convert speech data into text. Based on this text data, it performs sentiment analysis to infer the visitor's intent. The input is speech data, and the output is text data and the sentiment analysis results.

[0321] Step 4:

[0322] The server assesses the level of risk.

[0323] The server scores the visitor's risk level based on facial recognition and sentiment analysis results. This evaluation takes into account the visitor's past history and emotional state. The input is the facial recognition result and sentiment analysis result, and the output is the risk level score.

[0324] Step 5:

[0325] Notify the device and the user.

[0326] Based on the assessed risk score, the server either plays an audio message to the visitor via the terminal or sends a notification to the user's mobile device. The input is the risk score, and the output is the automated response on the terminal or the notification to the user.

[0327] Step 6:

[0328] The user selects the next action.

[0329] Users check the information conveyed via their smartphones and take action as needed, such as reporting to the police or notifying family members. The input is the notification information from the system, and the output is the specific action to be taken.

[0330] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0331] This invention combines a system that uses facial images and voice to assess visitor safety with an emotion engine that recognizes user emotions. This system functions in conjunction with the terminal, server, user, and emotion engine.

[0332] The device is installed at the entrance of the house and is equipped with a camera to capture images of visitors' faces and a microphone to record their voices. The acquired facial images and voice data are transmitted to a server in real time.

[0333] The server matches received facial images against a database to identify visitors, estimates their emotions through voice analysis, and assesses their risk level. The emotion engine also analyzes user voice and touch actions to recognize their emotions. This enables the presentation of information and responses tailored to the user's emotional state.

[0334] For example, if the emotion engine detects anxiety from the user's voice, the server can simplify the display of visitor information and provide additional reassuring information to alleviate the user's anxiety. Furthermore, it can recommend automated response options to ensure the user can respond comfortably.

[0335] As a concrete example, when a visitor arrives, the device retrieves the visitor's information, and the server assesses the level of risk. Once the assessment results are sent, the emotion engine recognizes the user's emotions. If the user is feeling anxious, the emotion engine instructs the system to display reassuring information to the user, presenting them with the message, "This visitor is safe." It can also recommend example responses that will help the user interact with the visitor in a friendly manner. This allows the user to interact with visitors with confidence.

[0336] By incorporating an emotion engine in this way, it becomes possible to go beyond simply assessing the risk level of visitors and instead provide detailed responses based on the user's emotions, thereby improving their sense of security.

[0337] The following describes the processing flow.

[0338] Step 1:

[0339] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0340] Step 2:

[0341] The device sends the acquired facial image data and audio data to the server. The server receives this data in real time.

[0342] Step 3:

[0343] The server analyzes the received facial images using a facial recognition algorithm, compares them with records in the database, and identifies the visitor. In particular, it checks for matches with past visit history and criminal lists.

[0344] Step 4:

[0345] The server analyzes the audio data using speech analysis technology. It estimates the speaker's emotions and intentions from the audio and analyzes the visitor's level of caution and trustworthiness.

[0346] Step 5:

[0347] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. This allows for an assessment of the visitor's safety.

[0348] Step 6:

[0349] The server sends instructions to the terminal based on the risk assessment. The terminal then executes an appropriate automated response to the visitor based on those instructions.

[0350] Step 7:

[0351] If a device is identified as a high-risk visitor, an alarm sounds to warn those nearby of the danger, and a danger notification is immediately sent from the server to the user.

[0352] Step 8:

[0353] The emotion engine analyzes the user's voice and touch actions to recognize the user's emotional state. Based on the user's emotions, it optimizes information presentation and response options.

[0354] Step 9:

[0355] Users can access the system via smartphones or computers to view visitor information and ratings. They can then use responses suggested by the emotion engine to remotely address situations or instruct police to be notified.

[0356] Step 10:

[0357] The server receives instructions from users and takes action, such as automatically reporting issues. It also adds evaluation results and visitor information to a database to record data for future reference.

[0358] (Example 2)

[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0360] Conventional visitor identification systems only assess the safety of visitors and lack consideration for the user's emotional state. Specifically, they did not provide guidance on how users should respond to visitors, making it difficult to respond appropriately. Furthermore, accurately assessing the visitor's emotions was difficult when determining their level of caution.

[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0362] In this invention, the server includes a device for acquiring images and audio to identify visitors, a device for analyzing the acquired images and audio to determine the visitor's level of alertness, and a device for detecting the user's emotional state and adjusting the information presented. This enables not only the assessment of the visitor's safety but also flexible responses and the provision of safety information in accordance with the user's emotional state.

[0363] "Image" refers to visual information used to visually record and analyze the faces and appearances of visitors.

[0364] "Audio" refers to auditory information used to record and analyze sounds emitted by visitors and users.

[0365] "Device" refers to a combination of hardware and software used to acquire, analyze, and display data.

[0366] "Alert level" refers to a measure of safety that is assessed based on the behavior and emotions of visitors.

[0367] "User emotional state" refers to the results of analyzing the emotional responses a user exhibits from their voice and behavior.

[0368] "Information presentation" refers to guidance and messages displayed to the user based on the analysis results.

[0369] This system identifies visitors, assesses their safety, and provides appropriate responses based on the user's emotional state. The system operates through the collaborative efforts of a terminal, a server, and an engine with emotional assessment capabilities.

[0370] Device features:

[0371] The device is installed at the entrance of the house and, when a visitor appears, captures their facial image with a camera and records their voice with a microphone. This data is transmitted to a server via secure communication. The device can be implemented, for example, using a camera module based on a Raspberry Pi and a high-sensitivity microphone for voice acquisition.

[0372] Server functions:

[0373] The server performs facial recognition by comparing received facial images with internal data. It can utilize facial recognition algorithms such as FaceNet or the OpenCV library. For audio data, it performs speech analysis using the Google Cloud Speech-to-Text API to estimate the visitor's emotions. This allows for a comprehensive evaluation of the visitor's level of caution. Based on this evaluation, information is provided to the user in real time.

[0374] Emotional engine function:

[0375] The user's emotional state is evaluated by an emotion engine based on voice tone and device touch operations. Real-time deep learning is utilized, such as with NVIDIA Jetson, to achieve emotion analysis. Based on the emotion engine's analysis results, information and responses tailored to the user's state are integrated.

[0376] Specific example:

[0377] When a visitor arrives at the front door, the device sends a facial image and audio to the server. The server determines the visitor is safe, and even if the emotion engine detects anxiety from the user's voice, it provides the user with reassuring information.

[0378] Example of a prompt:

[0379] "The visitor's facial image and audio data have been sent to the server. Please evaluate whether this visitor is safe and, based on that information, display a reassuring message to the user."

[0380] The above system enables visitor safety assessment and user support, providing users with a sense of security.

[0381] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0382] Step 1:

[0383] When the terminal detects a visitor's presence, it captures a facial image with its camera and records audio with its microphone. This data generates facial information as an image and audio information, which are then sent to the server in digital format. Specifically, a camera module using a Raspberry Pi captures an image, and an audio device captures clear audio.

[0384] Step 2:

[0385] The server processes the received facial image through a facial recognition algorithm and compares it with existing facial data within the system. The input is the transmitted facial image, and the output is the result of the facial recognition. This determines whether the visitor is a person who has been identified in the past. This process is performed using facial recognition technology such as FaceNet.

[0386] Step 3:

[0387] Simultaneously, the server inputs the audio data into a speech sentiment analysis tool to perform calculations that estimate the visitor's emotions. The input is the audio data, and the output is the estimated emotional state. The Google Cloud Speech-to-Text API and sentiment analysis libraries are used for speech-to-text conversion and emotion estimation.

[0388] Step 4:

[0389] The server comprehensively evaluates the visitor's level of caution based on the results of facial recognition and voice emotion estimation. The input is the facial recognition result and the voice emotion evaluation result, and the output is the caution level. This evaluation is a process that quantifies how much of a risk the visitor poses to the user.

[0390] Step 5:

[0391] The emotion engine analyzes the user's emotions. It analyzes the user's emotional state based on their voice and device interactions. Input is the user's voice or touch data, and output is the detected emotional state. Data processing using NVIDIA Jetson enables real-time estimation of the user's emotions.

[0392] Step 6:

[0393] The server generates information to guide the user based on the visitor's alert level and the user's emotional state. The input is the visitor's alert level and the user's emotional state, and the output is a message directed at the user. A generative AI model is used to provide specific messages and response examples as needed.

[0394] Each step works in conjunction with the others to create a system that supports visitor safety assessment and user interaction.

[0395] (Application Example 2)

[0396] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0397] In recent years, with the increase in visitors, there has been a growing need for an efficient system that can quickly and accurately assess visitor safety and provide users with a sense of security. Conventional systems are limited to providing information about visitors and do not provide appropriate information tailored to the user's emotional state, leaving the possibility of users experiencing tension or anxiety. This invention aims to solve such problems and improve the sense of security users experience when interacting with visitors.

[0398] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0399] In this invention, the server includes means for acquiring images and sounds to identify visitors, means for analyzing the acquired data to evaluate the visitor's safety, and means for adjusting the displayed content based on the user's emotional state and providing reassuring information. This makes it possible to provide users with immediate and appropriate information regarding the visitor's safety, thereby improving their sense of security.

[0400] A "visitor" is an individual whose facial image and voice are acquired by the system, and whose safety and emotional state are evaluated.

[0401] "Images" are visual data acquired to visually capture the characteristics of visitors.

[0402] "Acoustics" refers to auditory data acquired to capture visitors' voices and other auditory characteristics.

[0403] "Means" refer to the structure or process established by a system to achieve a specific function.

[0404] A "server" is a computer device that receives and analyzes acquired data and processes information in cooperation with other devices and systems.

[0405] A "user" is someone who receives visitor information through the system and provides instructions or responses to visitors.

[0406] "Safety" is the result of an assessment that visitors do not pose a potential risk based on the data collected.

[0407] The system of the present invention includes a sophisticated process for evaluating the safety of visitors and providing users with a sense of security. The terminal collects the visitor's facial image and voice and transmits them to the server. The server analyzes the received data and performs a safety evaluation. Specifically, it matches the facial image with a storage medium and analyzes the voice to identify emotions.

[0408] The server then uses a generative AI model to analyze the user's emotional state and, if necessary, generates information to provide reassurance. For example, it can display a message such as "This visitor is safe" to a user who is feeling uneasy about the visitor.

[0409] The software used includes machine learning libraries such as TensorFlow for facial image recognition and the Google Cloud Speech-to-Text API for speech data analysis. IBM Watson Tone Analyzer is also used to assess user emotions. This allows the system to instantly present information appropriate to the user's state.

[0410] As a concrete example, when a delivery person visits a home, the terminal captures the delivery person's face and voice. The server analyzes this data to confirm that the delivery person's face is registered in the database and detects a friendly tone of voice. If the user is feeling stressed, the system displays a message such as "The delivery person is safe and reliable" to enhance the user's sense of security.

[0411] An example of a prompt message might be, "Assess the visitor's safety and emotional state based on their facial image and voice, and display information to alleviate their anxiety." This process allows users to interact with visitors with confidence.

[0412] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0413] Step 1:

[0414] The device acquires facial images and audio from visitors. It uses raw data obtained from the camera and microphone as input. This data is recorded as image and audio data.

[0415] Step 2:

[0416] The terminal sends the acquired facial image and audio data to the server. The input is the data acquired in step 1, and the output is digital data converted into a format that can be received by the server.

[0417] Step 3:

[0418] The server identifies the visitor by comparing the received facial image with the storage medium. Using the facial image data as input, it performs a database search and outputs the visitor's ID and related information as a result.

[0419] Step 4:

[0420] The server converts audio data into text using the Google Cloud Speech-to-Text API and performs sentiment analysis based on that text data. The input is audio data, and the output is an evaluation result indicating the emotional state.

[0421] Step 5:

[0422] The server uses IBM Watson Tone Analyzer to analyze the user's emotional state. The input is a visitor safety assessment generated internally by the server, and the output determines the reassuring information that should be provided to the user.

[0423] Step 6:

[0424] The server sends data to the user's terminal to display visitor information and safety assessment results. The input is the results of steps 3 and 5, and the output is the message to be displayed and the recommended action.

[0425] Step 7:

[0426] The terminal displays information received from the server to the user and informs them of the system's final result. Input is data from the server, and output is information displayed on the terminal's screen. This allows the user to respond appropriately to visitors.

[0427] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0428] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0429] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0430] [Third Embodiment]

[0431] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0432] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0433] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0434] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0435] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0436] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0437] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0438] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0439] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0440] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0441] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0442] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0443] This invention provides a system that acquires and analyzes visitor information and assesses risk levels to provide appropriate responses to visitors. This system is operated by integrating terminals, servers, and users.

[0444] The terminal is installed at the building's entrance and is equipped with a camera and microphone to capture images of visitors' faces and record their voices. The terminal transmits this data to a server in real time.

[0445] The server uses facial recognition and voice analysis technologies to analyze data transmitted from the terminal. The server compares the visitor's face to known facial data in the database and checks for matches with visit history and criminal lists. Furthermore, the server infers emotions and intentions from the visitor's voice and scores the visitor's risk level.

[0446] Based on the risk level assessed by the server, the terminal automatically responds to visitors. Responses can be pre-set messages or voice messages via the intercom. The terminal can also sound an alarm if necessary.

[0447] Users can access the system in real time from their smartphones or computers via the internet. They can view visitor information and risk assessments provided by the server, and, if necessary, respond remotely or instruct the police to be notified. This allows users to respond appropriately to visitors even when they are away, ensuring the safety of their homes.

[0448] For example, when a visitor appears at the front door, the device automatically takes a picture of their face and records their voice. The server analyzes this data, and if it determines that the visitor is a legitimate delivery person with a history of no problems, the device plays an automated message such as, "Thank you for the delivery, please leave your package at the front door." On the other hand, if the server detects any suspicious behavior from the visitor, it issues an alarm to warn the user, and a notification is immediately sent to the user. After reviewing the situation, the user can, if necessary, report it to the police or share the details of the situation with their family.

[0449] Thus, the system of the present invention enables a quick and accurate response to visitors and serves as an effective means of enhancing the security function of a residence.

[0450] The following describes the processing flow.

[0451] Step 1:

[0452] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0453] Step 2:

[0454] The device transmits facial image data and audio data acquired by the device to the server in real time. The server receives this data.

[0455] Step 3:

[0456] The server analyzes facial images using a facial recognition algorithm and compares them with records in the database. In particular, it checks for matches with past visit history and criminal lists.

[0457] Step 4:

[0458] The server analyzes the audio data. It analyzes the characteristics of the voice and applies speech processing techniques to estimate the speaker's emotions and intentions.

[0459] Step 5:

[0460] The server integrates facial recognition results and voice analysis results to score the visitor's risk level. This provides a safety assessment of the visitor.

[0461] Step 6:

[0462] The server sends an appropriate automated response instruction to the terminal based on the risk assessment. The terminal receives this instruction and sends a response message to the visitor.

[0463] Step 7:

[0464] If the device determines that a visitor poses a high risk, it will sound an alarm to warn those nearby. Additionally, a danger notification will be immediately sent from the server to the user.

[0465] Step 8:

[0466] The user receives a notification, accesses it via the internet, and verifies visitor information and risk assessment. If necessary, remote responses or instructions to report to the police are provided.

[0467] Step 9:

[0468] The server receives instructions from the user and automatically notifies the police. It also updates the database and stores information to be used for future visitor detection.

[0469] (Example 1)

[0470] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0471] In recent years, the need for home security has been increasing. However, there is a problem in quickly and accurately understanding the intentions of visitors and responding appropriately, whether the homeowner is at home or away. Furthermore, existing security systems lack the ability to quantitatively assess the risk posed by visitors and provide appropriate real-time notifications to users. This invention aims to solve these problems.

[0472] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0473] In this invention, the server includes means for acquiring images and sounds using an imaging device to identify a visitor, means for evaluating the visitor's level of caution by processing the acquired images and sounds, and means for transmitting visitor information and caution level information to the user. This makes it possible to quickly analyze the visitor's intentions and level of caution and to immediately provide the user with necessary notifications.

[0474] An "imaging device" is a device used to acquire data such as images and sounds, and generally includes cameras and microphones.

[0475] An "image" is data that represents visual information such as a visitor's face or body.

[0476] "Acoustics" refers to data that represents auditory information, including visitors' voices and ambient sounds.

[0477] "Processing" refers to a series of actions taken to analyze acquired image and audio data and derive specific information.

[0478] "Alert level" is an indicator that quantifies or assigns numerical values ​​to evaluate the level of risk a visitor poses.

[0479] An "information recording device" is a device that stores data and lists of known visitors and verifies that information as needed.

[0480] A "user" is an individual or organization that operates this system and interacts with visitors.

[0481] "Notifications" refer to actions or functions that inform users about visitor information and the level of alertness.

[0482] "Sequential delivery" means providing information in real time by sending information continuously without any time gaps.

[0483] "Response instructions" refer to instructions or reactions that users give to visitors through the system.

[0484] This invention is a system that acquires a visitor's facial image and voice, analyzes this data to evaluate the visitor's level of caution, and provides an appropriate response to the visitor immediately. This system consists of a terminal, a server, and a user.

[0485] The terminal is installed at the building's entrance and used to capture images and audio of visitors. A high-resolution imaging device (camera) and a high-sensitivity audio input device (microphone) are used to quickly capture the visitor's facial and voice characteristics and transmit them as digital data to a server.

[0486] The server is the central device for processing the acquired data. Image data processing uses image analysis libraries such as OpenCV to identify visitors' faces by comparing them to data in the database. Furthermore, audio processing modules such as the Google Cloud Speech-to-Text API are used for speech analysis to estimate emotions and intentions from the visitor's voice. Based on these analysis results, the server scores the visitor's level of caution, enabling real-time decision-making.

[0487] Based on data processing by the server, if a visitor is deemed high-risk, the terminal will either play an automated response message or issue an alarm. This ensures a safe response without direct interaction.

[0488] Users can access the system via the internet to check visitor information and alert levels in real time. Using smart devices, users can receive notifications and respond appropriately no matter where they are. This allows users to maintain the security of their homes with peace of mind, even when they are away.

[0489] As a concrete example, when a visitor appears at the front door, the terminal takes a picture of the person with its camera and records their voice. The server analyzes this data, and if it recognizes, for example, the person as a delivery person who has visited the house before, the terminal will automatically respond with, "Thank you for the delivery, please leave the package at the front door." Furthermore, if there is anything suspicious about the visitor's behavior, the terminal will sound an alarm and immediately notify the user. An example of a prompt to the generating AI model in this case would be, "Please explain how the system works to analyze the visitor's facial image and voice data to assess the level of caution. Also, please give specific examples of responses."

[0490] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0491] Step 1:

[0492] The device captures a facial image with its camera and records audio with its microphone when a visitor arrives at the building entrance. The input consists of visual and audio information of the visitor's face, and the output is the storage of this data in a digital format. This process utilizes a high-resolution camera and a high-sensitivity microphone, and is designed to adapt to changes in the external environment.

[0493] Step 2:

[0494] The device transmits acquired facial images and audio data to the server in real time. The input is the digital data collected in step 1, and the output is data transmitted to the server in an encrypted format. This process uses a secure communication protocol to protect the confidentiality of the data.

[0495] Step 3:

[0496] The server processes the received facial image data through a facial recognition algorithm. The input data is the transmitted facial image, and the output is the result of matching it with a known facial database. Specifically, the OpenCV library is used here to determine whether the visitor is a registered person.

[0497] Step 4:

[0498] The server processes audio data using a speech analysis module. The input is the transmitted audio data, and the output is the analysis results of emotions and intentions extracted from the audio. The Google Cloud Speech-to-Text API is used to convert the audio to text, and then sentiment analysis is performed to evaluate the visitor's psychological state.

[0499] Step 5:

[0500] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. The input is the results of facial recognition and voice analysis, and the output is the visitor's risk level score. This score is determined based on criteria defined by the server and assesses the overall risk level of the visitor.

[0501] Step 6:

[0502] The terminal plays an automated response message or issues an alarm based on the alert level score received from the server. The input is the alert level score calculated in step 5, and the output is a voice message or alarm for the visitor. Specifically, if the visitor is friendly, a message such as "Welcome" is played, and if there is a reason to be cautious, an alarm sounds immediately.

[0503] (Application Example 1)

[0504] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0505] In modern society, ensuring the safety of visitors to homes and facilities is a critical issue. Furthermore, there is a need to quickly analyze visitor information and take appropriate action. However, existing security systems struggle to assess visitor risk in real time, making it difficult for users to effectively manage them remotely. Therefore, there is a need for the development of a system that accurately analyzes visitor risk and allows users to take immediate action.

[0506] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0507] In this invention, the server includes receiving means for identifying visitors, evaluation means for analyzing acquired data to assess the visitor's risk level, and communication means for notifying the user of the visitor's information and risk level assessment. This enables real-time assessment of the visitor's safety level, allowing the user to check the risk level via a mobile device and respond quickly.

[0508] "Receiving means" refers to a device for acquiring facial images and voices of visitors.

[0509] "Evaluation means" refers to a device or method for analyzing facial images and audio acquired by a receiving means to evaluate the level of risk a visitor poses.

[0510] "Communication means" refers to a device or method for notifying users of visitor information and risk assessments.

[0511] "Remote monitoring means" refers to a device or method that enables a user to check the safety level of a visitor in real time via a portable electronic device.

[0512] This invention relates to a security system that analyzes visitor information in real time to enhance the safety of a residence. A server receives facial images and voice data from a terminal installed at the building's entrance to identify visitors. Facial images are analyzed using the "face_recognition" library and compared against known facial data in a database. Voice data is converted to text using the "speech_recognition" library, and its emotions are analyzed.

[0513] Based on this data, the server assesses the visitor's risk level and notifies the terminal or the user's portable electronic device of the risk level via communication means. Users can check the visitor's safety level in real time via their smartphone or other mobile device and take prompt and appropriate action as needed.

[0514] For example, when a delivery person arrives, the server recognizes their face, and if it determines from past history that they are safe, it plays a message on the device such as, "Thank you for the delivery, please leave your package at the doorstep." In the unlikely event that the visitor is deemed suspicious, an alarm is issued and the user is immediately notified.

[0515] An example of a prompt message would be, "When a visitor arrives, we will capture their photo and audio to determine if they are known or suspicious. If they are safe, we will send a notification to your smartphone." In this way, users can effectively manage the security of their homes and facilities.

[0516] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0517] Step 1:

[0518] The device captures the visitor's facial image and voice.

[0519] The device uses a camera and microphone to capture facial images and audio data of visitors. This data is transmitted to the server in real time. The input is raw image and audio data, and the output is the transmission of data to the server.

[0520] Step 2:

[0521] The server analyzes the facial image data.

[0522] The server analyzes the received face image data using the "face_recognition" library. It compares it with known face data in the database to determine if there is a match. The input is face image data, and the output is the match result and identification information.

[0523] Step 3:

[0524] The server analyzes the audio data.

[0525] The server uses the "speech_recognition" library to convert speech data into text. Based on this text data, it performs sentiment analysis to infer the visitor's intent. The input is speech data, and the output is text data and the sentiment analysis results.

[0526] Step 4:

[0527] The server assesses the level of risk.

[0528] The server scores the visitor's risk level based on facial recognition and sentiment analysis results. This evaluation takes into account the visitor's past history and emotional state. The input is the facial recognition result and sentiment analysis result, and the output is the risk level score.

[0529] Step 5:

[0530] Notify the device and the user.

[0531] Based on the assessed risk score, the server either plays an audio message to the visitor via the terminal or sends a notification to the user's mobile device. The input is the risk score, and the output is the automated response on the terminal or the notification to the user.

[0532] Step 6:

[0533] The user selects the next action.

[0534] Users check the information conveyed via their smartphones and take action as needed, such as reporting to the police or notifying family members. The input is the notification information from the system, and the output is the specific action to be taken.

[0535] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0536] This invention combines a system that uses facial images and voice to assess visitor safety with an emotion engine that recognizes user emotions. This system functions in conjunction with the terminal, server, user, and emotion engine.

[0537] The device is installed at the entrance of the house and is equipped with a camera to capture images of visitors' faces and a microphone to record their voices. The acquired facial images and voice data are transmitted to a server in real time.

[0538] The server matches received facial images against a database to identify visitors, estimates their emotions through voice analysis, and assesses their risk level. The emotion engine also analyzes user voice and touch actions to recognize their emotions. This enables the presentation of information and responses tailored to the user's emotional state.

[0539] For example, if the emotion engine detects anxiety from the user's voice, the server can simplify the display of visitor information and provide additional reassuring information to alleviate the user's anxiety. Furthermore, it can recommend automated response options to ensure the user can respond comfortably.

[0540] As a concrete example, when a visitor arrives, the device retrieves the visitor's information, and the server assesses the level of risk. Once the assessment results are sent, the emotion engine recognizes the user's emotions. If the user is feeling anxious, the emotion engine instructs the system to display reassuring information to the user, presenting them with the message, "This visitor is safe." It can also recommend example responses that will help the user interact with the visitor in a friendly manner. This allows the user to interact with visitors with confidence.

[0541] By incorporating an emotion engine in this way, it becomes possible to go beyond simply assessing the risk level of visitors and instead provide detailed responses based on the user's emotions, thereby improving their sense of security.

[0542] The following describes the processing flow.

[0543] Step 1:

[0544] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0545] Step 2:

[0546] The device sends the acquired facial image data and audio data to the server. The server receives this data in real time.

[0547] Step 3:

[0548] The server analyzes the received facial images using a facial recognition algorithm, compares them with records in the database, and identifies the visitor. In particular, it checks for matches with past visit history and criminal lists.

[0549] Step 4:

[0550] The server analyzes the audio data using speech analysis technology. It estimates the speaker's emotions and intentions from the audio and analyzes the visitor's level of caution and trustworthiness.

[0551] Step 5:

[0552] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. This allows for an assessment of the visitor's safety.

[0553] Step 6:

[0554] The server sends instructions to the terminal based on the risk assessment. The terminal then executes an appropriate automated response to the visitor based on those instructions.

[0555] Step 7:

[0556] If a device is identified as a high-risk visitor, an alarm sounds to warn those nearby of the danger, and a danger notification is immediately sent from the server to the user.

[0557] Step 8:

[0558] The emotion engine analyzes the user's voice and touch actions to recognize the user's emotional state. Based on the user's emotions, it optimizes information presentation and response options.

[0559] Step 9:

[0560] Users can access the system via smartphones or computers to view visitor information and ratings. They can then use responses suggested by the emotion engine to remotely address situations or instruct police to be notified.

[0561] Step 10:

[0562] The server receives instructions from users and takes action, such as automatically reporting issues. It also adds evaluation results and visitor information to a database to record data for future reference.

[0563] (Example 2)

[0564] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0565] Conventional visitor identification systems only assess the safety of visitors and lack consideration for the user's emotional state. Specifically, they did not provide guidance on how users should respond to visitors, making it difficult to respond appropriately. Furthermore, accurately assessing the visitor's emotions was difficult when determining their level of caution.

[0566] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0567] In this invention, the server includes a device for acquiring images and audio to identify visitors, a device for analyzing the acquired images and audio to determine the visitor's level of alertness, and a device for detecting the user's emotional state and adjusting the information presented. This enables not only the assessment of the visitor's safety but also flexible responses and the provision of safety information in accordance with the user's emotional state.

[0568] "Image" refers to visual information used to visually record and analyze the faces and appearances of visitors.

[0569] "Audio" refers to auditory information used to record and analyze sounds emitted by visitors and users.

[0570] "Device" refers to a combination of hardware and software used to acquire, analyze, and display data.

[0571] "Alert level" refers to a measure of safety that is assessed based on the behavior and emotions of visitors.

[0572] "User emotional state" refers to the results of analyzing the emotional responses a user exhibits from their voice and behavior.

[0573] "Information presentation" refers to guidance and messages displayed to the user based on the analysis results.

[0574] This system identifies visitors, assesses their safety, and provides appropriate responses based on the user's emotional state. The system operates through the collaborative efforts of a terminal, a server, and an engine with emotional assessment capabilities.

[0575] Device features:

[0576] The device is installed at the entrance of the house and, when a visitor appears, captures their facial image with a camera and records their voice with a microphone. This data is transmitted to a server via secure communication. The device can be implemented, for example, using a camera module based on a Raspberry Pi and a high-sensitivity microphone for voice acquisition.

[0577] Server functions:

[0578] The server performs facial recognition by comparing received facial images with internal data. It can utilize facial recognition algorithms such as FaceNet or the OpenCV library. For audio data, it performs speech analysis using the Google Cloud Speech-to-Text API to estimate the visitor's emotions. This allows for a comprehensive evaluation of the visitor's level of caution. Based on this evaluation, information is provided to the user in real time.

[0579] Emotional engine function:

[0580] The user's emotional state is evaluated by an emotion engine based on voice tone and device touch operations. Real-time deep learning is utilized, such as with NVIDIA Jetson, to achieve emotion analysis. Based on the emotion engine's analysis results, information and responses tailored to the user's state are integrated.

[0581] Specific example:

[0582] When a visitor arrives at the front door, the device sends a facial image and audio to the server. The server determines the visitor is safe, and even if the emotion engine detects anxiety from the user's voice, it provides the user with reassuring information.

[0583] Example of a prompt:

[0584] "The visitor's facial image and audio data have been sent to the server. Please assess whether this visitor is safe and use that information to display a reassuring message to the user."

[0585] The above system enables visitor safety assessment and user support, providing users with a sense of security.

[0586] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0587] Step 1:

[0588] When the terminal detects a visitor's presence, it captures a facial image with its camera and records audio with its microphone. This data generates facial information as an image and audio information, which are then sent to the server in digital format. Specifically, a camera module using a Raspberry Pi captures an image, and an audio device captures clear audio.

[0589] Step 2:

[0590] The server processes the received facial image through a facial recognition algorithm and compares it with existing facial data within the system. The input is the transmitted facial image, and the output is the result of the facial recognition. This determines whether the visitor is a person who has been identified in the past. This process is performed using facial recognition technology such as FaceNet.

[0591] Step 3:

[0592] Simultaneously, the server inputs the audio data into a speech sentiment analysis tool to perform calculations that estimate the visitor's emotions. The input is the audio data, and the output is the estimated emotional state. The Google Cloud Speech-to-Text API and sentiment analysis libraries are used for speech-to-text conversion and emotion estimation.

[0593] Step 4:

[0594] The server comprehensively evaluates the visitor's level of caution based on the results of facial recognition and voice emotion estimation. The input is the facial recognition result and the voice emotion evaluation result, and the output is the caution level. This evaluation is a process that quantifies how much of a risk the visitor poses to the user.

[0595] Step 5:

[0596] The emotion engine analyzes the user's emotions. It analyzes the user's emotional state based on their voice and device interactions. Input is the user's voice or touch data, and output is the detected emotional state. Data processing using NVIDIA Jetson enables real-time estimation of the user's emotions.

[0597] Step 6:

[0598] The server generates information to guide the user based on the visitor's alert level and the user's emotional state. The input is the visitor's alert level and the user's emotional state, and the output is a message directed at the user. A generative AI model is used to provide specific messages and response examples as needed.

[0599] Each step works in conjunction with the others to create a system that supports visitor safety assessment and user interaction.

[0600] (Application Example 2)

[0601] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0602] In recent years, with the increase in visitors, there has been a growing need for an efficient system that can quickly and accurately assess visitor safety and provide users with a sense of security. Conventional systems are limited to providing information about visitors and do not provide appropriate information tailored to the user's emotional state, leaving the possibility of users experiencing tension or anxiety. This invention aims to solve such problems and improve the sense of security users experience when interacting with visitors.

[0603] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0604] In this invention, the server includes means for acquiring images and sounds to identify visitors, means for analyzing the acquired data to evaluate the visitor's safety, and means for adjusting the displayed content based on the user's emotional state and providing reassuring information. This makes it possible to provide users with immediate and appropriate information regarding the visitor's safety, thereby improving their sense of security.

[0605] A "visitor" is an individual whose facial image and voice are acquired by the system, and whose safety and emotional state are evaluated.

[0606] "Images" are visual data acquired to visually capture the characteristics of visitors.

[0607] "Acoustics" refers to auditory data acquired to capture visitors' voices and other auditory characteristics.

[0608] "Means" refer to the structure or process established by a system to achieve a specific function.

[0609] A "server" is a computer device that receives and analyzes acquired data and processes information in cooperation with other devices and systems.

[0610] A "user" is someone who receives visitor information through the system and provides instructions or responses to visitors.

[0611] "Safety" is the result of an assessment that visitors do not pose a potential risk based on the data collected.

[0612] The system of the present invention includes a sophisticated process for evaluating the safety of visitors and providing users with a sense of security. The terminal collects the visitor's facial image and voice and transmits them to the server. The server analyzes the received data and performs a safety evaluation. Specifically, it matches the facial image with a storage medium and analyzes the voice to identify emotions.

[0613] The server then uses a generative AI model to analyze the user's emotional state and, if necessary, generates information to provide reassurance. For example, it can display a message such as "This visitor is safe" to a user who is feeling uneasy about the visitor.

[0614] The software used includes machine learning libraries such as TensorFlow for facial image recognition and the Google Cloud Speech-to-Text API for speech data analysis. IBM Watson Tone Analyzer is also used to assess user emotions. This allows the system to instantly present information appropriate to the user's state.

[0615] As a concrete example, when a delivery person visits a home, the terminal captures the delivery person's face and voice. The server analyzes this data to confirm that the delivery person's face is registered in the database and detects a friendly tone of voice. If the user is feeling stressed, the system displays a message such as "The delivery person is safe and reliable" to enhance the user's sense of security.

[0616] An example of a prompt message might be, "Assess the visitor's safety and emotional state based on their facial image and voice, and display information to alleviate their anxiety." This process allows users to interact with visitors with confidence.

[0617] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0618] Step 1:

[0619] The device acquires facial images and audio from visitors. It uses raw data obtained from the camera and microphone as input. This data is recorded as image and audio data.

[0620] Step 2:

[0621] The terminal sends the acquired facial image and audio data to the server. The input is the data acquired in step 1, and the output is digital data converted into a format that can be received by the server.

[0622] Step 3:

[0623] The server identifies the visitor by comparing the received facial image with the storage medium. Using the facial image data as input, it performs a database search and outputs the visitor's ID and related information as a result.

[0624] Step 4:

[0625] The server converts audio data into text using the Google Cloud Speech-to-Text API and performs sentiment analysis based on that text data. The input is audio data, and the output is an evaluation result indicating the emotional state.

[0626] Step 5:

[0627] The server uses IBM Watson Tone Analyzer to analyze the user's emotional state. The input is a visitor safety assessment generated internally by the server, and the output determines the reassuring information that should be provided to the user.

[0628] Step 6:

[0629] The server sends data to the user's terminal to display visitor information and safety assessment results. The input is the results of steps 3 and 5, and the output is the message to be displayed and the recommended action.

[0630] Step 7:

[0631] The terminal displays information received from the server to the user and informs them of the system's final result. Input is data from the server, and output is information displayed on the terminal's screen. This allows the user to respond appropriately to visitors.

[0632] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0633] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0634] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0635] [Fourth Embodiment]

[0636] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0637] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0638] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0639] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0640] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0641] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0642] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0643] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0644] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0645] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0646] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0647] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0648] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0649] This invention provides a system that acquires and analyzes visitor information and assesses risk levels to provide appropriate responses to visitors. This system is operated by integrating terminals, servers, and users.

[0650] The terminal is installed at the building's entrance and is equipped with a camera and microphone to capture images of visitors' faces and record their voices. The terminal transmits this data to a server in real time.

[0651] The server uses facial recognition and voice analysis technologies to analyze data transmitted from the terminal. The server compares the visitor's face to known facial data in the database and checks for matches with visit history and criminal lists. Furthermore, the server infers emotions and intentions from the visitor's voice and scores the visitor's risk level.

[0652] Based on the risk level assessed by the server, the terminal automatically responds to visitors. Responses can be pre-set messages or voice messages via the intercom. The terminal can also sound an alarm if necessary.

[0653] Users can access the system in real time from their smartphones or computers via the internet. They can view visitor information and risk assessments provided by the server, and, if necessary, respond remotely or instruct the police to be notified. This allows users to respond appropriately to visitors even when they are away, ensuring the safety of their homes.

[0654] For example, when a visitor appears at the front door, the device automatically takes a picture of their face and records their voice. The server analyzes this data, and if it determines that the visitor is a legitimate delivery person with a history of no problems, the device plays an automated message such as, "Thank you for the delivery, please leave your package at the front door." On the other hand, if the server detects any suspicious behavior from the visitor, it issues an alarm to warn the user, and a notification is immediately sent to the user. After reviewing the situation, the user can, if necessary, report it to the police or share the details of the situation with their family.

[0655] Thus, the system of the present invention enables a quick and accurate response to visitors and serves as an effective means of enhancing the security function of a residence.

[0656] The following describes the processing flow.

[0657] Step 1:

[0658] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0659] Step 2:

[0660] The device transmits facial image data and audio data acquired by the device to the server in real time. The server receives this data.

[0661] Step 3:

[0662] The server analyzes facial images using a facial recognition algorithm and compares them with records in the database. In particular, it checks for matches with past visit history and criminal lists.

[0663] Step 4:

[0664] The server analyzes the audio data. It analyzes the characteristics of the voice and applies speech processing techniques to estimate the speaker's emotions and intentions.

[0665] Step 5:

[0666] The server integrates facial recognition results and voice analysis results to score the visitor's risk level. This provides a safety assessment of the visitor.

[0667] Step 6:

[0668] The server sends an appropriate automated response instruction to the terminal based on the risk assessment. The terminal receives this instruction and sends a response message to the visitor.

[0669] Step 7:

[0670] If the device determines that a visitor poses a high risk, it will sound an alarm to warn those nearby. Additionally, a danger notification will be immediately sent from the server to the user.

[0671] Step 8:

[0672] The user receives a notification, accesses it via the internet, and verifies visitor information and risk assessment. If necessary, remote responses or instructions to report to the police are provided.

[0673] Step 9:

[0674] The server receives instructions from the user and automatically notifies the police. It also updates the database and stores information to be used for future visitor detection.

[0675] (Example 1)

[0676] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] In recent years, the need for home security has been increasing. However, there is a problem in quickly and accurately understanding the intentions of visitors and responding appropriately, whether the homeowner is at home or away. Furthermore, existing security systems lack the ability to quantitatively assess the risk posed by visitors and provide appropriate real-time notifications to users. This invention aims to solve these problems.

[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0679] In this invention, the server includes means for acquiring images and sounds using an imaging device to identify a visitor, means for evaluating the visitor's level of caution by processing the acquired images and sounds, and means for transmitting visitor information and caution level information to the user. This makes it possible to quickly analyze the visitor's intentions and level of caution and to immediately provide the user with necessary notifications.

[0680] An "imaging device" is a device used to acquire data such as images and sounds, and generally includes cameras and microphones.

[0681] An "image" is data that represents visual information such as a visitor's face or body.

[0682] "Acoustics" refers to data that represents auditory information, including visitors' voices and ambient sounds.

[0683] "Processing" refers to a series of actions taken to analyze acquired image and audio data and derive specific information.

[0684] "Alert level" is an indicator that quantifies or assigns numerical values ​​to evaluate the level of risk a visitor poses.

[0685] An "information recording device" is a device that stores data and lists of known visitors and verifies that information as needed.

[0686] A "user" is an individual or organization that operates this system and interacts with visitors.

[0687] "Notifications" refer to actions or functions that inform users about visitor information and the level of alertness.

[0688] "Sequential delivery" means providing information in real time by sending information continuously without any time gaps.

[0689] "Response instructions" refer to instructions or reactions that users give to visitors through the system.

[0690] This invention is a system that acquires a visitor's facial image and voice, analyzes this data to evaluate the visitor's level of caution, and provides an appropriate response to the visitor immediately. This system consists of a terminal, a server, and a user.

[0691] The terminal is installed at the building's entrance and used to capture images and audio of visitors. A high-resolution imaging device (camera) and a high-sensitivity audio input device (microphone) are used to quickly capture the visitor's facial and voice characteristics and transmit them as digital data to a server.

[0692] The server is the central device for processing the acquired data. Image data processing uses image analysis libraries such as OpenCV to identify visitors' faces by comparing them to data in the database. Furthermore, audio processing modules such as the Google Cloud Speech-to-Text API are used for speech analysis to estimate emotions and intentions from the visitor's voice. Based on these analysis results, the server scores the visitor's level of caution, enabling real-time decision-making.

[0693] Based on data processing by the server, if a visitor is deemed high-risk, the terminal will either play an automated response message or issue an alarm. This ensures a safe response without direct interaction.

[0694] Users can access the system via the internet to check visitor information and alert levels in real time. Using smart devices, users can receive notifications and respond appropriately no matter where they are. This allows users to maintain the security of their homes with peace of mind, even when they are away.

[0695] As a concrete example, when a visitor appears at the front door, the terminal takes a picture of the person with its camera and records their voice. The server analyzes this data, and if it recognizes, for example, the person as a delivery person who has visited the house before, the terminal will automatically respond with, "Thank you for the delivery, please leave the package at the front door." Furthermore, if there is anything suspicious about the visitor's behavior, the terminal will sound an alarm and immediately notify the user. An example of a prompt to the generating AI model in this case would be, "Please explain how the system works to analyze the visitor's facial image and voice data to assess the level of caution. Also, please give specific examples of responses."

[0696] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0697] Step 1:

[0698] The device captures a facial image with its camera and records audio with its microphone when a visitor arrives at the building entrance. The input consists of visual and audio information of the visitor's face, and the output is the storage of this data in a digital format. This process utilizes a high-resolution camera and a high-sensitivity microphone, and is designed to adapt to changes in the external environment.

[0699] Step 2:

[0700] The device transmits acquired facial images and audio data to the server in real time. The input is the digital data collected in step 1, and the output is data transmitted to the server in an encrypted format. This process uses a secure communication protocol to protect the confidentiality of the data.

[0701] Step 3:

[0702] The server processes the received facial image data through a facial recognition algorithm. The input data is the transmitted facial image, and the output is the result of matching it with a known facial database. Specifically, the OpenCV library is used here to determine whether the visitor is a registered person.

[0703] Step 4:

[0704] The server processes audio data using a speech analysis module. The input is the transmitted audio data, and the output is the analysis results of emotions and intentions extracted from the audio. The Google Cloud Speech-to-Text API is used to convert the audio to text, and then sentiment analysis is performed to evaluate the visitor's psychological state.

[0705] Step 5:

[0706] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. The input is the results of facial recognition and voice analysis, and the output is the visitor's risk level score. This score is determined based on criteria defined by the server and assesses the overall risk level of the visitor.

[0707] Step 6:

[0708] The terminal plays an automated response message or issues an alarm based on the alert level score received from the server. The input is the alert level score calculated in step 5, and the output is a voice message or alarm for the visitor. Specifically, if the visitor is friendly, a message such as "Welcome" is played, and if there is a reason to be cautious, an alarm sounds immediately.

[0709] (Application Example 1)

[0710] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0711] In modern society, ensuring the safety of visitors to homes and facilities is a critical issue. Furthermore, there is a need to quickly analyze visitor information and take appropriate action. However, existing security systems struggle to assess visitor risk in real time, making it difficult for users to effectively manage them remotely. Therefore, there is a need for the development of a system that accurately analyzes visitor risk and allows users to take immediate action.

[0712] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0713] In this invention, the server includes receiving means for identifying visitors, evaluation means for analyzing acquired data to assess the visitor's risk level, and communication means for notifying the user of the visitor's information and risk level assessment. This enables real-time assessment of the visitor's safety level, allowing the user to check the risk level via a mobile device and respond quickly.

[0714] "Receiving means" refers to a device for acquiring facial images and voices of visitors.

[0715] "Evaluation means" refers to a device or method for analyzing facial images and audio acquired by a receiving means to evaluate the level of risk a visitor poses.

[0716] "Communication means" refers to a device or method for notifying users of visitor information and risk assessments.

[0717] "Remote monitoring means" refers to a device or method that enables a user to check the safety level of a visitor in real time via a portable electronic device.

[0718] This invention relates to a security system that analyzes visitor information in real time to enhance the safety of a residence. A server receives facial images and voice data from a terminal installed at the building's entrance to identify visitors. Facial images are analyzed using the "face_recognition" library and compared against known facial data in a database. Voice data is converted to text using the "speech_recognition" library, and its emotions are analyzed.

[0719] Based on this data, the server assesses the visitor's risk level and notifies the terminal or the user's portable electronic device of the risk level via communication means. Users can check the visitor's safety level in real time via their smartphone or other mobile device and take prompt and appropriate action as needed.

[0720] For example, when a delivery person arrives, the server recognizes their face, and if it determines from past history that they are safe, it plays a message on the device such as, "Thank you for the delivery, please leave your package at the doorstep." In the unlikely event that the visitor is deemed suspicious, an alarm is issued and the user is immediately notified.

[0721] An example of a prompt message would be, "When a visitor arrives, we will capture their photo and audio to determine if they are known or suspicious. If they are safe, we will send a notification to your smartphone." In this way, users can effectively manage the security of their homes and facilities.

[0722] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0723] Step 1:

[0724] The device captures the visitor's facial image and voice.

[0725] The device uses a camera and microphone to capture facial images and audio data of visitors. This data is transmitted to the server in real time. The input is raw image and audio data, and the output is the transmission of data to the server.

[0726] Step 2:

[0727] The server analyzes the facial image data.

[0728] The server analyzes the received face image data using the "face_recognition" library. It compares it with known face data in the database to determine if there is a match. The input is face image data, and the output is the match result and identification information.

[0729] Step 3:

[0730] The server analyzes the audio data.

[0731] The server uses the "speech_recognition" library to convert speech data into text. Based on this text data, it performs sentiment analysis to infer the visitor's intent. The input is speech data, and the output is text data and the sentiment analysis results.

[0732] Step 4:

[0733] The server assesses the level of risk.

[0734] The server scores the visitor's risk level based on facial recognition and sentiment analysis results. This evaluation takes into account the visitor's past history and emotional state. The input is the facial recognition result and sentiment analysis result, and the output is the risk level score.

[0735] Step 5:

[0736] Notify the device and the user.

[0737] Based on the assessed risk score, the server either plays an audio message to the visitor via the terminal or sends a notification to the user's mobile device. The input is the risk score, and the output is the automated response on the terminal or the notification to the user.

[0738] Step 6:

[0739] The user selects the next action.

[0740] Users check the information conveyed via their smartphones and take action as needed, such as reporting to the police or notifying family members. The input is the notification information from the system, and the output is the specific action to be taken.

[0741] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0742] This invention combines a system that uses facial images and voice to assess visitor safety with an emotion engine that recognizes user emotions. This system functions in conjunction with the terminal, server, user, and emotion engine.

[0743] The device is installed at the entrance of the house and is equipped with a camera to capture images of visitors' faces and a microphone to record their voices. The acquired facial images and voice data are transmitted to a server in real time.

[0744] The server matches received facial images against a database to identify visitors, estimates their emotions through voice analysis, and assesses their risk level. The emotion engine also analyzes user voice and touch actions to recognize their emotions. This enables the presentation of information and responses tailored to the user's emotional state.

[0745] For example, if the emotion engine detects anxiety from the user's voice, the server can simplify the display of visitor information and provide additional reassuring information to alleviate the user's anxiety. Furthermore, it can recommend automated response options to ensure the user can respond comfortably.

[0746] As a concrete example, when a visitor arrives, the device retrieves the visitor's information, and the server assesses the level of risk. Once the assessment results are sent, the emotion engine recognizes the user's emotions. If the user is feeling anxious, the emotion engine instructs the system to display reassuring information to the user, presenting them with the message, "This visitor is safe." It can also recommend example responses that will help the user interact with the visitor in a friendly manner. This allows the user to interact with visitors with confidence.

[0747] By incorporating an emotion engine in this way, it becomes possible to go beyond simply assessing the risk level of visitors and instead provide detailed responses based on the user's emotions, thereby improving their sense of security.

[0748] The following describes the processing flow.

[0749] Step 1:

[0750] The device uses a sensor installed at the entrance to detect the approach of a visitor. It captures the visitor's face with a camera and records their voice using a microphone.

[0751] Step 2:

[0752] The device sends the acquired facial image data and audio data to the server. The server receives this data in real time.

[0753] Step 3:

[0754] The server analyzes the received facial images using a facial recognition algorithm, compares them with records in the database, and identifies the visitor. In particular, it checks for matches with past visit history and criminal lists.

[0755] Step 4:

[0756] The server analyzes the audio data using speech analysis technology. It estimates the speaker's emotions and intentions from the audio and analyzes the visitor's level of caution and trustworthiness.

[0757] Step 5:

[0758] The server integrates the results of facial recognition and voice analysis to score the visitor's risk level. This allows for an assessment of the visitor's safety.

[0759] Step 6:

[0760] The server sends instructions to the terminal based on the risk assessment. The terminal then executes an appropriate automated response to the visitor based on those instructions.

[0761] Step 7:

[0762] If a device is identified as a high-risk visitor, an alarm sounds to warn those nearby of the danger, and a danger notification is immediately sent from the server to the user.

[0763] Step 8:

[0764] The emotion engine analyzes the user's voice and touch actions to recognize the user's emotional state. Based on the user's emotions, it optimizes information presentation and response options.

[0765] Step 9:

[0766] Users can access the system via smartphones or computers to view visitor information and ratings. They can then use responses suggested by the emotion engine to remotely address situations or instruct police to be notified.

[0767] Step 10:

[0768] The server receives instructions from users and takes action, such as automatically reporting issues. It also adds evaluation results and visitor information to a database to record data for future reference.

[0769] (Example 2)

[0770] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0771] Conventional visitor identification systems only assess the safety of visitors and lack consideration for the user's emotional state. Specifically, they did not provide guidance on how users should respond to visitors, making it difficult to respond appropriately. Furthermore, accurately assessing the visitor's emotions was difficult when determining their level of caution.

[0772] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0773] In this invention, the server includes a device for acquiring images and audio to identify visitors, a device for analyzing the acquired images and audio to determine the visitor's level of alertness, and a device for detecting the user's emotional state and adjusting the information presented. This enables not only the assessment of the visitor's safety but also flexible responses and the provision of safety information in accordance with the user's emotional state.

[0774] "Image" refers to visual information used to visually record and analyze the faces and appearances of visitors.

[0775] "Audio" refers to auditory information used to record and analyze sounds emitted by visitors and users.

[0776] "Device" refers to a combination of hardware and software used to acquire, analyze, and display data.

[0777] "Alert level" refers to a measure of safety that is assessed based on the behavior and emotions of visitors.

[0778] "User emotional state" refers to the results of analyzing the emotional responses a user exhibits from their voice and behavior.

[0779] "Information presentation" refers to guidance and messages displayed to the user based on the analysis results.

[0780] This system identifies visitors, assesses their safety, and provides appropriate responses based on the user's emotional state. The system operates through the collaborative efforts of a terminal, a server, and an engine with emotional assessment capabilities.

[0781] Device features:

[0782] The device is installed at the entrance of the house and, when a visitor appears, captures their facial image with a camera and records their voice with a microphone. This data is transmitted to a server via secure communication. The device can be implemented, for example, using a camera module based on a Raspberry Pi and a high-sensitivity microphone for voice acquisition.

[0783] Server functions:

[0784] The server performs facial recognition by comparing received facial images with internal data. It can utilize facial recognition algorithms such as FaceNet or the OpenCV library. For audio data, it performs speech analysis using the Google Cloud Speech-to-Text API to estimate the visitor's emotions. This allows for a comprehensive evaluation of the visitor's level of caution. Based on this evaluation, information is provided to the user in real time.

[0785] Emotional engine function:

[0786] The user's emotional state is evaluated by an emotion engine based on voice tone and device touch operations. Real-time deep learning is utilized, such as with NVIDIA Jetson, to achieve emotion analysis. Based on the emotion engine's analysis results, information and responses tailored to the user's state are integrated.

[0787] Specific example:

[0788] When a visitor arrives at the front door, the device sends a facial image and audio to the server. The server determines the visitor is safe, and even if the emotion engine detects anxiety from the user's voice, it provides the user with reassuring information.

[0789] Example of a prompt:

[0790] "The visitor's facial image and audio data have been sent to the server. Please assess whether this visitor is safe and use that information to display a reassuring message to the user."

[0791] The above system enables visitor safety assessment and user support, providing users with a sense of security.

[0792] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0793] Step 1:

[0794] When the terminal detects a visitor's presence, it captures a facial image with its camera and records audio with its microphone. This data generates facial information as an image and audio information, which are then sent to the server in digital format. Specifically, a camera module using a Raspberry Pi captures an image, and an audio device captures clear audio.

[0795] Step 2:

[0796] The server processes the received facial image through a facial recognition algorithm and compares it with existing facial data within the system. The input is the transmitted facial image, and the output is the result of the facial recognition. This determines whether the visitor is a person who has been identified in the past. This process is performed using facial recognition technology such as FaceNet.

[0797] Step 3:

[0798] Simultaneously, the server inputs the audio data into a speech sentiment analysis tool to perform calculations that estimate the visitor's emotions. The input is the audio data, and the output is the estimated emotional state. The Google Cloud Speech-to-Text API and sentiment analysis libraries are used for speech-to-text conversion and emotion estimation.

[0799] Step 4:

[0800] The server comprehensively evaluates the visitor's level of caution based on the results of facial recognition and voice emotion estimation. The input is the facial recognition result and the voice emotion evaluation result, and the output is the caution level. This evaluation is a process that quantifies how much of a risk the visitor poses to the user.

[0801] Step 5:

[0802] The emotion engine analyzes the user's emotions. It analyzes the user's emotional state based on their voice and device interactions. Input is the user's voice or touch data, and output is the detected emotional state. Data processing using NVIDIA Jetson enables real-time estimation of the user's emotions.

[0803] Step 6:

[0804] The server generates information to guide the user based on the visitor's alert level and the user's emotional state. The input is the visitor's alert level and the user's emotional state, and the output is a message directed at the user. A generative AI model is used to provide specific messages and response examples as needed.

[0805] Each step works in conjunction with the others to create a system that supports visitor safety assessment and user interaction.

[0806] (Application Example 2)

[0807] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0808] In recent years, with the increase in visitors, there has been a growing need for an efficient system that can quickly and accurately assess visitor safety and provide users with a sense of security. Conventional systems are limited to providing information about visitors and do not provide appropriate information tailored to the user's emotional state, leaving the possibility of users experiencing tension or anxiety. This invention aims to solve such problems and improve the sense of security users experience when interacting with visitors.

[0809] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0810] In this invention, the server includes means for acquiring images and sounds to identify visitors, means for analyzing the acquired data to evaluate the visitor's safety, and means for adjusting the displayed content based on the user's emotional state and providing reassuring information. This makes it possible to provide users with immediate and appropriate information regarding the visitor's safety, thereby improving their sense of security.

[0811] A "visitor" is an individual whose facial image and voice are acquired by the system, and whose safety and emotional state are evaluated.

[0812] "Images" are visual data acquired to visually capture the characteristics of visitors.

[0813] "Acoustics" refers to auditory data acquired to capture visitors' voices and other auditory characteristics.

[0814] "Means" refer to the structure or process established by a system to achieve a specific function.

[0815] A "server" is a computer device that receives and analyzes acquired data and processes information in cooperation with other devices and systems.

[0816] A "user" is someone who receives visitor information through the system and provides instructions or responses to visitors.

[0817] "Safety" is the result of an assessment that visitors do not pose a potential risk based on the data collected.

[0818] The system of the present invention includes a sophisticated process for evaluating the safety of visitors and providing users with a sense of security. The terminal collects the visitor's facial image and voice and transmits them to the server. The server analyzes the received data and performs a safety evaluation. Specifically, it matches the facial image with a storage medium and analyzes the voice to identify emotions.

[0819] The server then uses a generative AI model to analyze the user's emotional state and, if necessary, generates information to provide reassurance. For example, it can display a message such as "This visitor is safe" to a user who is feeling uneasy about the visitor.

[0820] The software used includes machine learning libraries such as TensorFlow for facial image recognition and the Google Cloud Speech-to-Text API for speech data analysis. IBM Watson Tone Analyzer is also used to assess user emotions. This allows the system to instantly present information appropriate to the user's state.

[0821] As a concrete example, when a delivery person visits a home, the terminal captures the delivery person's face and voice. The server analyzes this data to confirm that the delivery person's face is registered in the database and detects a friendly tone of voice. If the user is feeling stressed, the system displays a message such as "The delivery person is safe and reliable" to enhance the user's sense of security.

[0822] An example of a prompt message might be, "Assess the visitor's safety and emotional state based on their facial image and voice, and display information to alleviate their anxiety." This process allows users to interact with visitors with confidence.

[0823] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0824] Step 1:

[0825] The device acquires facial images and audio from visitors. It uses raw data obtained from the camera and microphone as input. This data is recorded as image and audio data.

[0826] Step 2:

[0827] The terminal sends the acquired facial image and audio data to the server. The input is the data acquired in step 1, and the output is digital data converted into a format that can be received by the server.

[0828] Step 3:

[0829] The server identifies the visitor by comparing the received facial image with the storage medium. Using the facial image data as input, it performs a database search and outputs the visitor's ID and related information as a result.

[0830] Step 4:

[0831] The server converts audio data into text using the Google Cloud Speech-to-Text API and performs sentiment analysis based on that text data. The input is audio data, and the output is an evaluation result indicating the emotional state.

[0832] Step 5:

[0833] The server uses IBM Watson Tone Analyzer to analyze the user's emotional state. The input is a visitor safety assessment generated internally by the server, and the output determines the reassuring information that should be provided to the user.

[0834] Step 6:

[0835] The server sends data to the user's terminal to display visitor information and safety assessment results. The input is the results of steps 3 and 5, and the output is the message to be displayed and the recommended action.

[0836] Step 7:

[0837] The terminal displays information received from the server to the user and informs them of the system's final result. Input is data from the server, and output is information displayed on the terminal's screen. This allows the user to respond appropriately to visitors.

[0838] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0839] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0840] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0841] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0842] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0843] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0844] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0845] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0846] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0847] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0848] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0849] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0850] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0851] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0852] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0853] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0854] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0855] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0856] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0857] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0858] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0859] The following is further disclosed regarding the embodiments described above.

[0860] (Claim 1)

[0861] A means for acquiring facial images and audio to identify visitors,

[0862] An evaluation method that analyzes acquired facial images and audio to assess the risk level of visitors,

[0863] A response system that automatically responds to visitors based on the assessed risk level,

[0864] A notification method that informs the user of visitor information and risk assessment,

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, characterized in that the evaluation means identifies visitors by matching their facial images with a database and evaluates their level of caution by analyzing the emotions in their voice.

[0868] (Claim 3)

[0869] The system according to claim 1, characterized in that the notification means provides a function to deliver visitor information and risk level to the user in real time and to allow the user to remotely issue response instructions.

[0870] "Example 1"

[0871] (Claim 1)

[0872] Means for acquiring images and sounds using an imaging device to identify visitors,

[0873] A means for evaluating the level of caution of visitors by processing acquired images and sounds,

[0874] A means of automatically generating a response to visitors based on the assessed level of alertness,

[0875] A means of transmitting visitor information and alert level information to users,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, characterized in that the evaluation means can identify visitors by comparing images with an information recording device and evaluate the level of alertness by analyzing the emotional impact of sounds.

[0879] (Claim 3)

[0880] The system according to claim 1, characterized in that the notification means has the function of sequentially distributing visitor information and alert levels to the user, and allowing the user to issue response instructions remotely.

[0881] "Application Example 1"

[0882] (Claim 1)

[0883] A receiving means that acquires facial images and audio in order to identify the person visiting,

[0884] An evaluation method that analyzes acquired facial images and audio to assess the risk level of visitors,

[0885] A response system that automatically responds to visitors based on the assessed risk level,

[0886] A communication method for notifying users of visitor information and risk assessment,

[0887] A remote monitoring system that uses facial images and audio to assess the safety level of visitors in real time, and allows users to check the level of risk via their mobile devices.

[0888] A system that includes this.

[0889] (Claim 2)

[0890] The system according to claim 1, characterized in that the evaluation means identifies a visitor by matching a facial image with an information recording device and evaluates the level of caution by analyzing the emotion of the voice.

[0891] (Claim 3)

[0892] The system according to claim 1, characterized in that the communication means has the function of delivering visitor information and risk level to the user in real time and allowing the user to remotely give response instructions, and includes displaying notifications on a portable electronic device.

[0893] "Example 2 of combining an emotion engine"

[0894] (Claim 1)

[0895] A device that acquires images and audio to identify visitors,

[0896] A device that analyzes acquired images and audio to determine the visitor's level of alertness,

[0897] A device that detects the user's emotional state and adjusts the information presented accordingly,

[0898] A device that displays visitor information and alert levels to users in real time,

[0899] ...

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, characterized in that it identifies visitors by comparing images and audio with internal records and evaluates the alert level by analyzing the emotional content of the sound.

[0903] (Claim 3)

[0904] The system according to claim 1, characterized in that it includes a function to transmit information and alert levels to the user according to the user's emotional state, and to allow the user to respond remotely.

[0905] "Application example 2 when combining with an emotional engine"

[0906] (Claim 1)

[0907] Means for acquiring images and sounds to identify visitors,

[0908] A means for analyzing acquired images and sounds to evaluate visitor safety,

[0909] A means of automatically displaying information to visitors based on the assessed safety,

[0910] A means of adjusting the displayed content according to the user's emotional state and providing information to reassure them,

[0911] A means of notifying users of visitor information and safety ratings,

[0912] A system that includes this.

[0913] (Claim 2)

[0914] The system according to claim 1, characterized in that the evaluation means includes a function to identify visitors by matching images with a storage medium, evaluate the level of alertness by analyzing the emotional nature of sounds, and provide additional reassurance information.

[0915] (Claim 3)

[0916] The system according to claim 1, characterized in that the notification means has a function to immediately transmit visitor information and security to the user and to allow the user to remotely instruct appropriate measures. [Explanation of Symbols]

[0917] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A receiving means that acquires facial images and audio in order to identify the person visiting, An evaluation method that analyzes acquired facial images and audio to assess the risk level of visitors, A response system that automatically responds to visitors based on the assessed risk level, A communication method for notifying users of visitor information and risk assessment, A remote monitoring system that uses facial images and audio to assess the safety level of visitors in real time, and allows users to check the level of risk via their mobile devices. A system that includes this.

2. The system according to claim 1, characterized in that the evaluation means identifies a visitor by matching a facial image with an information recording device and evaluates the level of caution by analyzing the emotion of the voice.

3. The system according to claim 1, characterized in that the communication means has the function of delivering visitor information and risk level to the user in real time and allowing the user to remotely issue response instructions, and includes displaying notifications on a portable electronic device.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A