System
The system uses a camera and microphone with AI analysis to identify and prevent crimes by evaluating visitor attributes and emotions, providing real-time security for vulnerable individuals.
Patent Information
- Application Number
- JP2024119050
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional security systems struggle to accurately identify suspicious visitors, particularly in cases of special fraud and fraud targeting elderly individuals, due to inadequate visitor identification through screens, increasing the risk of fraudsters making contact.
A system utilizing a camera and microphone to capture visitor video and audio, analyzed by a generative artificial intelligence model to extract attributes, calculate a score, and determine suspiciousness, with notification to users and authorities if necessary.
Enables real-time, accurate identification and prevention of crimes by automatically determining visitor suspiciousness, ensuring user safety and prompt action, especially for vulnerable groups like the elderly.
Smart Images

Figure 2026017989000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, crimes such as special fraud, deposit fraud, and "it's me" fraud still claim many victims, with elderly people being particularly prone to being targeted. Conventional security systems have difficulty identifying visitors through a screen, which increases the opportunities for fraudsters to make contact. A new system is needed to address these issues and prevent crime before it happens. [Means for solving the problem]
[0005] The present invention solves the above problem with a system including a means for capturing video of visitors with a camera, a means for capturing audio of the visitors with a microphone, an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes, a means for evaluating the visitor's attributes based on the extracted attributes and calculating a score, a means for determining whether the visitor is suspicious based on the calculated score, and a means for notifying the user and sending a warning to the police if the visitor is determined to be suspicious. This allows users, including elderly people, to automatically determine whether a visitor is suspicious and take appropriate action.
[0006] A "camera" is a device for capturing images of visitors.
[0007] A "microphone" is a device for capturing the voice of visitors.
[0008] "Video" refers to moving or still images captured by a camera.
[0009] "Audio" refers to the voices of visitors and surrounding sounds picked up by the microphone.
[0010] A "generative artificial intelligence model" is a machine learning model that analyzes video and audio data to extract visitor attributes.
[0011] "Analysis means" refers to the process of using a generative artificial intelligence model to analyze video and audio data and extract visitor attributes.
[0012] "Attributes" are characteristics such as a visitor's age, gender, whether they are a familiar face, or any anomalies in the background sounds.
[0013] A "score" is a numerical evaluation of a visitor based on the extracted attributes.
[0014] The "means of determination" is the process of determining whether a visitor is suspicious or not based on the calculated score.
[0015] "User" refers to an individual or household who uses this system to implement security measures.
[0016] A "warning" is the act of notifying the police or related agencies that a suspicious person is visiting.
[0017] "Notification" is the act of notifying the user when a suspicious person visits. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system for preventing special fraud and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and aims to detect visitors who may be committing fraud.
[0040] System Configuration
[0041] The present invention comprises the following main components:
[0042] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0043] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0044] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0045] 4. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates a score. The score is higher for elderly people, unfamiliar faces, and unnatural background sounds.
[0046] 5. The judgment means judges whether the visitor is suspicious or not based on the calculated score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0047] 6. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[0048] System Operation
[0049] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. The server then evaluates the extracted attributes using a score calculation means to calculate a score. Based on this score, the server determines whether the visitor is suspicious.
[0050] As a specific example, the following scenario can be considered.
[0051] Example scenario:
[0052] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0053] 2. The video shows a middle-aged man and the audio is mixed with unnatural noise. The video and audio data are automatically sent to an analysis device.
[0054] 3. The server analyzes the data, estimates the visitor's age to be 45, and verifies that the person has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background.
[0055] 4. Based on these attributes, the server calculates a score. For example, a high score can be set if the person's age is within a certain range, if they are not registered as an acquaintance, or if unnatural noise is detected.
[0056] 5. The visitor is determined to be suspicious based on the score calculated by the server.
[0057] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0058] By using such a system, users can protect themselves from crime, and it can provide an effective crime prevention measure, especially for those who are more likely to become victims of crime, such as the elderly.
[0059] The processing flow will be explained below.
[0060] Step 1:
[0061] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[0062] Step 2:
[0063] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[0064] Step 3:
[0065] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[0066] Step 4:
[0067] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0068] Step 5:
[0069] The server calculates a score for each visitor based on the extracted attributes. For example, an elderly person will receive a higher score, a familiar face will receive a lower score, and the presence of unnatural background noise will also result in a higher score.
[0070] Step 6:
[0071] The server uses the calculated score to determine whether the visitor is suspicious. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[0072] Step 7:
[0073] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[0074] Step 8:
[0075] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[0076] Step 9:
[0077] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[0078] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[0079] Example 1
[0080] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0081] In recent years, there has been an increase in special frauds and other crimes, with the elderly and other socially vulnerable people being particularly vulnerable. Conventional crime prevention measures monitor visitor information in real time, but it is often difficult to accurately identify suspicious individuals. Furthermore, ensuring the security of captured data and linking it to a prompt notification system for suspicious individuals have been issues.
[0082] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0083] In this invention, the server includes a means for capturing video of visitors using a camera, a means for capturing audio of visitors using a microphone, and a means for transmitting the captured and analyzed data using a secure protocol. This allows visitor information to be captured in real time, securely transmitted, and analyzed, making it possible to accurately evaluate visitor attributes and identify suspicious individuals. This also ensures data security and allows for prompt notification of suspicious individuals.
[0084] A "camera" is a device that captures video of visitors.
[0085] A "microphone" is a device that captures the visitor's voice.
[0086] A "generative artificial intelligence model" is an artificial intelligence model that analyzes captured video and audio data and extracts visitor attributes.
[0087] The "analysis means" is a means for analyzing the captured video and audio data to extract visitor attributes.
[0088] "Attributes" refers to characteristics such as the visitor's age, gender, whether they are a known person, and any abnormalities in background noise.
[0089] The "score calculation means" is a means for evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[0090] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[0091] The "notification means" is a means for notifying the user and the police if a person is determined to be a suspicious person.
[0092] A "secure protocol" is a communication method that ensures security in data transmission.
[0093] A "timestamp" is time information added to captured data.
[0094] A "data analysis module" is a piece of software functionality for extracting attributes from captured video and audio data.
[0095] A "score calculation module" is a part of the software function that calculates a score based on visitor attributes.
[0096] A "judgment module" is a part of the software function that judges whether a visitor is suspicious based on the calculated score.
[0097] This invention is a system for preventing special fraud and other crimes committed by visitors. The system is configured using the following hardware and software:
[0098] Hardware
[0099] 1. Camera: This is a device built into the intercom that captures video of visitors. When a visitor stands in front of the intercom, the camera automatically starts recording.
[0100] 2. Microphone: This is a device built into the intercom that captures the visitor's voice. When a visitor speaks, the microphone records the voice.
[0101] software
[0102] 1. Generative AI model: An AI model that analyzes captured video and audio data and extracts visitor attributes.
[0103] 2. Analysis method: A method for analyzing captured video and audio data to extract visitor attributes.
[0104] 3. Score calculation means: A means for evaluating the visitor's attributes based on the extracted attributes and calculating a score.
[0105] 4. Judgment method: A method for judging whether a visitor is suspicious or not based on the calculated score.
[0106] 5. Notification means: A means for notifying the user and the police based on the judgment results.
[0107] 6. Secure Protocol: A communication method that ensures security in data transmission.
[0108] 7. Data Analysis Module: This is the part of the software that functions to extract attributes from the captured video and audio data.
[0109] 8. Score calculation module: A part of the software function that calculates a score based on visitor attributes.
[0110] 9. Judgment module: This is the part of the software function that judges whether a visitor is suspicious or not based on the calculated score.
[0111] The system works as follows: First, the terminal captures the visitor's video and audio in real time using the intercom's built-in camera and microphone. Next, the captured data is sent to the server using a secure protocol (e.g., SSL / TLS). The server then analyzes the received video and audio data using a generative artificial intelligence model to extract attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. These attributes are evaluated using a score calculation means to calculate a score.
[0112] As a specific example, when a visitor presses the button on the intercom, the device activates the camera and microphone to capture the visitor's video and audio. The video shows a middle-aged man, and the audio is mixed with unnatural noise. The captured data is immediately sent to the server using a secure protocol, where it is analyzed. The analysis results indicate that the visitor is estimated to be 45 years old and that he has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background. Based on this information, the server calculates a high score and determines the visitor to be suspicious. Finally, a notification method is activated, sending an alert to the user's smartphone and also notifying the police.
[0113] Prompt Sentence Examples
[0114] "The video of a middle-aged man captured by an intercom camera contains unnatural noise. We will analyze this visitor to determine whether he is suspicious."
[0115] This system allows users to monitor visitor information in real time to ensure safety, and also provides effective crime prevention measures for vulnerable groups such as the elderly.
[0116] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0117] Step 1:
[0118] The device uses the intercom's camera and microphone to capture video and audio of visitors in real time. The camera automatically captures video when the visitor stands in front of the intercom, and the microphone records audio when the visitor speaks. The captured video and audio data are then time-stamped.
[0119] Input: Visitor video and audio
[0120] Output: Time-stamped video and audio data
[0121] How it works: When a visitor presses the button on the intercom, the camera and microphone automatically activate and record video and audio.
[0122] Step 2:
[0123] The device transmits the captured, time-stamped video and audio data to the server using a secure protocol (e.g., SSL / TLS), where the data is buffered and transmitted in real time.
[0124] Input: Time-stamped video and audio data
[0125] Output: Securely transmitted video and audio data
[0126] What happens: The device buffers the captured data and prepares to send it to the server over a secure channel. The transmission begins immediately.
[0127] Step 3:
[0128] The server uses a data analysis module to analyze the video and audio data it receives. First, video analysis involves facial recognition and age and gender estimation, while audio analysis involves the presence or absence of background noise and the characteristics of the audio.
[0129] Input: Securely transmitted video and audio data
[0130] Output: Visitor demographic data (e.g., age, gender, presence or absence of background noise)
[0131] How it works: The server analyzes the video data frame by frame to extract facial features, and simultaneously performs spectrogram analysis of the audio data to check for abnormal noise.
[0132] Step 4:
[0133] The server uses a score calculation module to calculate a score based on attributes extracted from the analysis results (e.g., the visitor's age, gender, whether they are a known person, whether there is background noise, etc.). The score is weighted for each attribute, and the overall likelihood of a suspicious person is evaluated.
[0134] Input: Visitor attribute data
[0135] Output: The calculated score
[0136] What happens: The server runs a program that estimates the visitor's age to be 45, determines their gender as male, checks for the presence of background noise, and calculates a score based on these attributes.
[0137] Step 5:
[0138] The server determines whether the visitor is suspicious based on the calculated score. If the visitor is determined to be suspicious, the server proceeds to the next notification process.
[0139] Input: Calculated score
[0140] Output: Result of suspicious person judgment (e.g., suspicious person / not suspicious person)
[0141] Specific operation: If the score exceeds a certain threshold, the server executes logic to determine the person as suspicious.
[0142] Step 6:
[0143] The notification mechanism will be activated, sending an alert to the user's smartphone and, if necessary, notifying the police, allowing the user to quickly understand the situation and take action.
[0144] Input: Suspicious person detection result
[0145] Output: Notification to user and police
[0146] Specific operation: The server sends a warning to the user's smartphone via SMS or app notification that there is a "suspicious person" and also notifies the police using an API.
[0147] (Application example 1)
[0148] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0149] Conventional visitor reception systems are good at collecting visitor information, but lack the ability to effectively analyze the collected information and quickly determine whether a visitor is suspicious. Furthermore, they lack the means to issue real-time warnings or reports, making it difficult to prevent crimes or take early action. The present invention aims to solve these problems by providing a system that prevents crimes committed by visitors and ensures the safety of users.
[0150] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0151] In this invention, the server includes means for capturing video of visitors with a camera, means for capturing audio of visitors with a microphone, analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes, means for evaluating the visitor attributes based on the extracted attributes and calculating a score, means for determining whether the visitor is a suspicious person based on the calculated score, means for notifying the user and sending a warning to the police if the visitor is determined to be suspicious, means for linking the intercom and smart device to transfer visitor information in real time, and means for analyzing visitor attributes in real time and calculating a score. This allows for quick and accurate analysis of visitor information, enabling the detection of suspicious people and immediate response.
[0152] A "camera" is a device that captures video of visitors.
[0153] A "microphone" is a device that captures the visitor's voice.
[0154] A "generative artificial intelligence model" is an analytical method that analyzes captured video and audio to extract visitor attributes.
[0155] The "score calculation means" is a means having a function of evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[0156] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[0157] "Notification means" refers to the means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[0158] "Means for linking an intercom with a smart device" refers to a means for linking an intercom with a smart device such as a smartphone to transfer visitor information in real time.
[0159] The "real-time analysis means" is a means for analyzing visitor attributes in real time and calculating scores.
[0160] The system for implementing the present invention captures video and audio of visitors in real time and analyzes them using a generative artificial intelligence model to detect suspicious individuals and issue immediate warnings. This system uses the following hardware and software:
[0161] Hardware and Software Configuration
[0162] Hardware
[0163] Camera: A device that captures video of visitors, usually built into the intercom.
[0164] Microphone: A device that captures the visitor's voice and is built into the intercom, just like a camera.
[0165] Smart device: Usually a smartphone, which works in conjunction with the intercom to notify the user of visitor information.
[0166] software
[0167] Generative artificial intelligence model: An AI model that analyzes captured video and audio to extract visitor attributes (e.g., age, gender, background noise), for example using a deep learning framework such as Keras.
[0168] Score calculation algorithm: An algorithm for calculating the score based on the extracted attributes. It is implemented using a programming language such as Python.
[0169] Notification systems: Systems that send alerts to users when a visitor is deemed suspicious. This includes email sending services (such as smtplib) and mobile notification systems.
[0170] Process Overview
[0171] 1. The device uses a camera and microphone to capture video and audio of visitors in real time, and the captured data is immediately transmitted to an analytical tool that includes a generative artificial intelligence model.
[0172] 2. The server analyzes the video and audio data using a generative artificial intelligence model, which extracts attributes such as the visitor's age, gender, and the presence or absence of background noise.
[0173] 3. The server uses a scoring algorithm to calculate the visitor's score based on the extracted attributes, specifically, unknown visitors and those with unnatural noise are given a higher score.
[0174] 4. Based on the calculated score, the server determines whether the visitor is suspicious.
[0175] 5. If a score exceeds the threshold specified by the user, the system will identify the person as suspicious and immediately send a warning to the user's smart device. If requested, the system will also automatically notify the police.
[0176] Specific examples
[0177] For example, if a middle-aged man is seen on camera as a visitor and unnatural noise is detected in the background, the generative AI model will estimate his age to be 45 and detect background noise from the audio data. Based on this information, the scoring means will assign a high score and determine that the visitor is suspicious. The user will then receive a real-time notification saying, "A suspicious visitor has been detected!"
[0178] An example of an input prompt for a generative AI model is as follows:
[0179] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[0180] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0181] Step 1:
[0182] The device uses a camera and microphone to capture video and audio of visitors in real time.
[0183] Input: Real-time video data, real-time audio data
[0184] Data processing: Video and audio data is acquired and converted into digital format.
[0185] Output: Digital video data, digital audio data
[0186] Step 2:
[0187] The terminal transfers the captured video and audio data to the server.
[0188] Input: Digital video data, digital audio data
[0189] Data processing: Packetizing data and transferring it over the network
[0190] Output: Video and audio data transferred to the server
[0191] Step 3:
[0192] The server uses a generative artificial intelligence model to analyze the captured video and audio data.
[0193] Input: Video data, audio data
[0194] Data processing: Data analysis using generative AI models (extracting attributes such as age, gender, and background noise)
[0195] Output: Visitor demographic data (age, gender, presence of background noise, etc.)
[0196] Step 4:
[0197] The server uses a scoring algorithm to calculate a score for the visitor based on the extracted attributes.
[0198] Input: Visitor attribute data
[0199] Data processing: Score calculation algorithm (adds score according to attributes)
[0200] Output: Visitor score
[0201] Step 5:
[0202] The server determines whether the visitor is suspicious based on the calculated score.
[0203] Input: Visitor score
[0204] Data processing: Comparing scores with set thresholds
[0205] Output: Suspicious person determination result (suspicious person or non-suspicious person)
[0206] Step 6:
[0207] If a person is determined to be suspicious, the server will send a warning to the user's smart device and, if necessary, notify the police.
[0208] Input: Suspicious person judgment result
[0209] Data processing: generating and sending notification messages (contacting users and the police)
[0210] Output: Warning notice to user, report to police
[0211] Considering the following specific example, an example prompt would be:
[0212] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[0213] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0214] This invention is a system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[0215] System Configuration
[0216] The present invention comprises the following main components:
[0217] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0218] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0219] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0220] 4. The emotion engine is a means of recognizing the user's emotions, and uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions.
[0221] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[0222] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0223] 7. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[0224] System Operation
[0225] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[0226] As a specific example, the following scenario can be considered.
[0227] Example scenario:
[0228] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0229] 2. The video shows a middle-aged woman, and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice. In front of the intercom, the user's face takes on an anxious expression and her voice becomes trembling.
[0230] 3. The server analyzes the data, estimates the visitor's age to be 35, and confirms that the person has not previously been registered as an acquaintance. Furthermore, audio analysis detects no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[0231] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[0232] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[0233] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0234] In this way, users can be more effectively protected from specialized fraud and other crimes by using a crime prevention system that combines emotion analysis, and the system can provide more advanced security measures to a wider range of people, including the elderly.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[0238] Step 2:
[0239] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[0240] Step 3:
[0241] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[0242] Step 4:
[0243] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0244] Step 5:
[0245] The device also captures the user's facial expressions and tone of voice to recognize their emotions. An emotion engine is used to analyze the user's emotions and identify emotions such as anxiety or fear.
[0246] Step 6:
[0247] The server calculates a score based on the extracted visitor attributes and the user's perceived emotions. For example, if an elderly person or unnatural noise is detected, the score will be higher. The score will also increase if the user shows signs of anxiety or fear.
[0248] Step 7:
[0249] The server determines whether the visitor is suspicious based on the calculated total score. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[0250] Step 8:
[0251] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[0252] Step 9:
[0253] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[0254] Step 10:
[0255] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[0256] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[0257] Example 2
[0258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0259] There is a need to quickly and accurately determine whether a visitor is suspicious and ensure the safety of users. Conventional systems only analyze visitor attributes and are unable to integrate information including user emotions. As a result, there is a problem that the accuracy of detecting suspicious individuals is low and users' anxiety cannot be completely alleviated.
[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0261] In this invention, the server includes a means for capturing video of the visitor with a camera, a means for capturing the visitor's voice with a microphone, and a means for capturing the user's facial expression and tone of voice and analyzing the user's emotions, thereby making it possible to determine suspicious individuals by taking into account not only the visitor's attributes but also the user's emotions.
[0262] A "camera" is a device for capturing images.
[0263] A "microphone" is a device for capturing sound.
[0264] A "generative artificial intelligence model" is an artificial intelligence technology that analyzes specific features and patterns from input data and outputs the results.
[0265] "Analysis means" refers to a device or software for analyzing the captured video and audio data and extracting visitor attributes.
[0266] "Means for capturing a user's facial expression and tone of voice and analyzing the user's emotions" refers to a device or software that uses a camera and microphone to capture a user's facial expression and tone of voice and, based on that, identifies the user's emotions.
[0267] The "means for calculating a score" is a device or software for evaluating visitor attributes based on the extracted visitor attributes and the analyzed user sentiment.
[0268] The "means for determining" is a device or software for determining whether a visitor is suspicious or not based on the calculated score.
[0269] The "notification means" is a device or system that notifies the user when a person is determined to be suspicious and also sends a warning to the police.
[0270] This invention is a security system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[0271] System Configuration
[0272] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0273] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0274] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data, which extracts attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. The generative artificial intelligence model used is a commonly used AI model (such as OpenAI's GPT-4).
[0275] 4. The emotion engine is a means of recognizing the user's emotions, and it analyzes the user's facial expressions and tone of voice using a camera and microphone. Emotion engines that can be used include Amazon's AWS Rekognition and Google Cloud Vision API.
[0276] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[0277] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0278] 7. If a person is determined to be suspicious, the notification method will notify the user and also send an alert to the police. This method allows the user to take appropriate action quickly. Services such as Twilio and Pushbullet can be used as notification methods.
[0279] System Operation
[0280] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[0281] Specific examples
[0282] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0283] 2. The video shows a middle-aged man and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice.
[0284] 3. The server analyzes the data, estimates the visitor's age to be 45, and confirms that this person has not previously been registered as an acquaintance. Furthermore, audio analysis detects that there is no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[0285] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[0286] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[0287] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0288] Examples of prompts:
[0289] Analyze the following data:
[0290] Visitor video data, age, gender, and whether they are known faces
[0291] Visitor voice data and whether there are any abnormal sounds in the background
[0292] Analysis results of user's facial expressions and tone of voice
[0293] The analysis methods are as follows:
[0294] 1. Identifying visitor attributes based on video data
[0295] 2. Analyze background sounds based on audio data
[0296] 3. Analyze user emotions
[0297] Calculate your overall score based on the analysis results.
[0298] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0299] Step 1:
[0300] Data Capture
[0301] The device captures the visitor's video and audio in real time. Specifically, the intercom captures video with a built-in camera and audio with a built-in microphone. The input is the visitor's real-time video and audio, and is used to output video and audio data.
[0302] Specific behavior:
[0303] The device's camera will then begin working, capturing high-resolution footage of the visitor.
[0304] The device's microphone will begin working and will record the visitor's speech and background sounds.
[0305] The captured video and audio data is temporarily stored in the device's memory.
[0306] Step 2:
[0307] Data transmission
[0308] The device sends the captured video and audio data to the server. The input is the captured video and audio data, which is then encrypted and sent to the server for output.
[0309] Specific behavior:
[0310] The terminal encrypts the video and audio data and transmits it to a server via a communications network.
[0311] Check whether the data was sent successfully and proceed to the next step.
[0312] Step 3:
[0313] Visitor attribute analysis
[0314] The server analyzes the received video data using a generative artificial intelligence model. The input is the video and audio data received from the device, and the output is the visitor's attribute data, such as age, gender, and whether or not the face is familiar.
[0315] Specific behavior:
[0316] The server inputs video data into the generative artificial intelligence model using prompt sentences and acquires attributes such as age and gender.
[0317] The server analyzes the audio data and determines whether there are any suspicious noises in the background sound.
[0318] The analysis results are stored in a database on the server.
[0319] Step 4:
[0320] User sentiment analysis
[0321] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and sends them to the server. The input is the user's facial expression data and tone of voice data, which is then used by the emotion engine to analyze the user's emotions and output them.
[0322] Specific behavior:
[0323] The device's camera captures the user's facial expressions in high resolution.
[0324] The device's microphone records the user's voice.
[0325] The captured data is sent to a server and analyzed by an emotion engine.
[0326] The analysis results, which represent the user's emotions (e.g., anxiety, fear, relief), are output to the server.
[0327] Step 5:
[0328] Calculation of overall score
[0329] The server calculates an overall score using the score calculation means based on the data obtained from the analysis means and emotion engine. The input is visitor attribute data and user emotion data, which are integrated to calculate and output an overall score.
[0330] Specific behavior:
[0331] The server combines visitor attribute data (e.g., age, gender, whether the face is familiar) with user emotion data (e.g., anxiety).
[0332] A score calculation means is used to quantify the overall score.
[0333] The calculated overall score is stored in a database on the server.
[0334] Step 6:
[0335] judgement
[0336] The server determines whether the visitor is suspicious or not based on the calculated total score. The input is the total score, and the output is the result of the determination whether the visitor is suspicious or not.
[0337] Specific behavior:
[0338] The server compares the overall score to a threshold.
[0339] If the score exceeds the threshold, the visitor is deemed suspicious.
[0340] The judgment results are stored in the server's database.
[0341] Step 7:
[0342] Notification and Response
[0343] If the visitor is determined to be suspicious, the device automatically responds to the visitor saying "We are unable to serve you," and the server sends a warning to the user's smartphone and even notifies the police. The input is the judgment result, and the output is the execution of the notification and report.
[0344] Specific behavior:
[0345] The device will send a voice notification to the visitor via the intercom saying "We are unavailable."
[0346] The server uses a service such as Twilio or Pushbullet to send an alert message to the user's smartphone.
[0347] The server notifies the police using a preset notification means.
[0348] (Application example 2)
[0349] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0350] In modern society, special frauds and other criminal activities are on the rise, posing a serious problem especially for the elderly and people living alone. Therefore, there is a need for advanced security systems that can efficiently analyze visitor identities and user emotions in real time to prevent criminal activities. However, conventional systems lack a comprehensive judgment function that takes into account not only visitor attributes but also user emotions, resulting in low fraud detection accuracy.
[0351] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0352] In this invention, the server includes: means for capturing video of visitors with a camera; means for capturing audio of visitors with a microphone; analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes; means for analyzing the user's video and audio to recognize the user's emotions; means for evaluating the visitor's attributes based on the extracted attributes and the recognized user emotions and calculating a score; means for determining whether the visitor is a suspicious person based on the calculated score; and means for notifying the user and sending a warning to the police if the visitor is determined to be a suspicious person. This realizes a comprehensive security system that integrates visitor identity analysis and user emotion recognition, making it possible to prevent specialized frauds and other criminal acts.
[0353] A "camera" is a device that captures video of visitors.
[0354] A "microphone" is a device that captures the visitor's voice.
[0355] A "generative artificial intelligence model" is a type of artificial intelligence that analyzes captured video and audio to extract visitor attributes.
[0356] "Analysis means" refers to means for analyzing captured video and audio, including generative artificial intelligence models.
[0357] A "user" is a person who uses this system.
[0358] The "means for recognizing emotions" is a means for identifying the emotions of a user by analyzing the user's video and audio.
[0359] The "means for evaluating attributes and calculating a score" is a means for evaluating attributes of a visitor based on the extracted attributes and the recognized user sentiment, and calculating a score.
[0360] The "means for determining whether a visitor is suspicious" is a means for determining whether a visitor is suspicious based on the calculated score.
[0361] The "means for notifying and sending a warning" is a means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[0362] This invention presents a form of advanced security system that analyzes visitor identity and user sentiment. The system consists of the following main components:
[0363] Hardware:
[0364] The system's terminal is a device equipped with a camera and microphone. This device functions as an intercom and captures real-time video and audio information of visitors, allowing the terminal to accurately record their actions and comments.
[0365] software:
[0366] The system applies a generative artificial intelligence model (AI model) and an emotion recognition engine that analyzes captured video and audio data to extract visitor attributes (age, gender, known or unknown) and user emotions (anxiety, relief, fear, etc.).
[0367] Analysis method:
[0368] Using AI models, the video and audio captured on the device are analyzed to extract visitor attributes.
[0369] Emotion recognition means:
[0370] The camera and microphone are used to analyze the user's facial expressions and tone of voice, and the emotion recognition engine identifies the user's emotions.
[0371] Data processing and calculation:
[0372] The server calculates a total score using the score calculation means based on the visitor's attributes extracted by the analysis means and the user's emotions identified by the emotion recognition means. This total score is used to determine whether the visitor is suspicious. Once a determination is made, the notification means is activated and a warning is sent to the user and the police.
[0373] Examples:
[0374] When a visitor standing at the front door operates the intercom, the camera automatically captures video and the microphone captures audio. This data is analyzed by an AI model to determine the visitor's age, gender, and whether they are a familiar face. Meanwhile, the user's camera video and audio are analyzed by an emotion recognition engine to identify the user's emotions. Based on this, the server calculates an overall score, integrating the visitor's attributes and the user's emotions to determine whether they are suspicious.
[0375] Example prompt sentence:
[0376] "Capture video of visitors and analyze their attributes."
[0377] "Estimate the visitor's age and gender and assess user sentiment."
[0378] "Use a comprehensive score to determine whether a visitor is suspicious."
[0379] In this way, the system provides comprehensive security that integrates visitor identity analysis and user emotion recognition, making it possible to effectively protect users from specialized fraud and other criminal activities.
[0380] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0381] Step 1:
[0382] The device uses a camera and microphone to capture the visitor's video and audio. The input is the visitor's real-time video and audio, and the output is the captured digital data. This digital data is temporarily stored for subsequent analysis processes.
[0383] Step 2:
[0384] The terminal sends the captured video and audio data to the analysis means. The input is the captured video and audio digital data, and the output is the visitor's attribute data (e.g., age, gender, whether the visitor is a known person) as the analysis result. The terminal analyzes the data using a generative artificial intelligence model and extracts the attributes.
[0385] Step 3:
[0386] The terminal captures the user's video and audio. This step is performed between the time the visitor operates the intercom and the time the user answers. The input is the user's real-time video and audio, and the output is the captured digital data.
[0387] Step 4:
[0388] The device sends the captured user video and audio data to the emotion recognition means. The input is the captured user video and audio digital data, and the output is the user's emotional data (e.g., anxiety, relief, fear, etc.) as an analysis result. The emotion recognition engine analyzes this data to identify the user's emotions.
[0389] Step 5:
[0390] The server integrates the output data of the analysis means and emotion recognition means and calculates an overall score using the score calculation means. The input is the visitor's attribute data and the user's emotion data, and the output is an overall score. The server performs data calculations based on this data and evaluates whether the visitor is suspicious or not.
[0391] Step 6:
[0392] The server determines whether the visitor is suspicious based on the total score. The input is the total score, and the output is the result of the judgment whether the visitor is suspicious. Based on this judgment result, a notification method is prepared.
[0393] Step 7:
[0394] If the server determines that a person is suspicious, it immediately notifies the user and sends a warning to the police. The input is the suspicious person determination result, and the output is a notification message to the user and a warning message to the police. The server automatically generates and sends the message.
[0395] In this way, the system analyzes the identity of visitors and the user's emotions in real time, enabling it to identify suspicious individuals and take prompt action.
[0396] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0397] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0398] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0399] [Second embodiment]
[0400] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0401] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0402] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0403] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0404] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0406] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0407] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0408] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0409] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0410] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0411] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0412] This invention is a system for preventing special fraud and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and aims to detect visitors who may be committing fraud.
[0413] System Configuration
[0414] The present invention comprises the following main components:
[0415] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0416] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0417] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0418] 4. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates a score. The score is higher for elderly people, unfamiliar faces, and unnatural background sounds.
[0419] 5. The judgment means judges whether the visitor is suspicious or not based on the calculated score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0420] 6. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[0421] System Operation
[0422] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. The server then evaluates the extracted attributes using a score calculation means to calculate a score. Based on this score, the server determines whether the visitor is suspicious.
[0423] As a specific example, the following scenario can be considered.
[0424] Example scenario:
[0425] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0426] 2. The video shows a middle-aged man and the audio is mixed with unnatural noise. The video and audio data are automatically sent to an analysis device.
[0427] 3. The server analyzes the data, estimates the visitor's age to be 45, and verifies that the person has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background.
[0428] 4. Based on these attributes, the server calculates a score. For example, a high score can be set if the person's age is within a certain range, if they are not registered as an acquaintance, or if unnatural noise is detected.
[0429] 5. The visitor is determined to be suspicious based on the score calculated by the server.
[0430] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0431] By using such a system, users can protect themselves from crime, and it can provide an effective crime prevention measure, especially for those who are more likely to become victims of crime, such as the elderly.
[0432] The processing flow will be explained below.
[0433] Step 1:
[0434] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[0435] Step 2:
[0436] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[0437] Step 3:
[0438] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[0439] Step 4:
[0440] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0441] Step 5:
[0442] The server calculates a score for each visitor based on the extracted attributes. For example, an elderly person will receive a higher score, a familiar face will receive a lower score, and the presence of unnatural background noise will also result in a higher score.
[0443] Step 6:
[0444] The server uses the calculated score to determine whether the visitor is suspicious. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[0445] Step 7:
[0446] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[0447] Step 8:
[0448] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[0449] Step 9:
[0450] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[0451] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[0452] Example 1
[0453] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0454] In recent years, there has been an increase in special frauds and other crimes, with the elderly and other socially vulnerable people being particularly vulnerable. Conventional crime prevention measures monitor visitor information in real time, but it is often difficult to accurately identify suspicious individuals. Furthermore, ensuring the security of captured data and linking it to a prompt notification system for suspicious individuals have been issues.
[0455] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0456] In this invention, the server includes a means for capturing video of visitors using a camera, a means for capturing audio of visitors using a microphone, and a means for transmitting the captured and analyzed data using a secure protocol. This allows visitor information to be captured in real time, securely transmitted, and analyzed, making it possible to accurately evaluate visitor attributes and identify suspicious individuals. This also ensures data security and allows for prompt notification of suspicious individuals.
[0457] A "camera" is a device that captures video of visitors.
[0458] A "microphone" is a device that captures the visitor's voice.
[0459] A "generative artificial intelligence model" is an artificial intelligence model that analyzes captured video and audio data and extracts visitor attributes.
[0460] The "analysis means" is a means for analyzing the captured video and audio data to extract visitor attributes.
[0461] "Attributes" refers to characteristics such as the visitor's age, gender, whether they are a known person, and any abnormalities in background noise.
[0462] The "score calculation means" is a means for evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[0463] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[0464] The "notification means" is a means for notifying the user and the police if a person is determined to be a suspicious person.
[0465] A "secure protocol" is a communication method that ensures security in data transmission.
[0466] A "timestamp" is time information added to captured data.
[0467] A "data analysis module" is a piece of software functionality for extracting attributes from captured video and audio data.
[0468] A "score calculation module" is a part of the software function that calculates a score based on visitor attributes.
[0469] A "judgment module" is a part of the software function that judges whether a visitor is suspicious based on the calculated score.
[0470] This invention is a system for preventing special fraud and other crimes committed by visitors. The system is configured using the following hardware and software:
[0471] Hardware
[0472] 1. Camera: This is a device built into the intercom that captures video of visitors. When a visitor stands in front of the intercom, the camera automatically starts recording.
[0473] 2. Microphone: This is a device built into the intercom that captures the visitor's voice. When a visitor speaks, the microphone records the voice.
[0474] software
[0475] 1. Generative AI model: An AI model that analyzes captured video and audio data and extracts visitor attributes.
[0476] 2. Analysis method: A method for analyzing captured video and audio data to extract visitor attributes.
[0477] 3. Score calculation means: A means for evaluating the visitor's attributes based on the extracted attributes and calculating a score.
[0478] 4. Judgment method: A method for judging whether a visitor is suspicious or not based on the calculated score.
[0479] 5. Notification means: A means for notifying the user and the police based on the judgment results.
[0480] 6. Secure Protocol: A communication method that ensures security in data transmission.
[0481] 7. Data Analysis Module: This is the part of the software that functions to extract attributes from the captured video and audio data.
[0482] 8. Score calculation module: A part of the software function that calculates a score based on visitor attributes.
[0483] 9. Judgment module: This is the part of the software function that judges whether a visitor is suspicious or not based on the calculated score.
[0484] The system works as follows: First, the terminal captures the visitor's video and audio in real time using the intercom's built-in camera and microphone. Next, the captured data is sent to the server using a secure protocol (e.g., SSL / TLS). The server then analyzes the received video and audio data using a generative artificial intelligence model to extract attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. These attributes are evaluated using a score calculation means to calculate a score.
[0485] As a specific example, when a visitor presses the button on the intercom, the device activates the camera and microphone to capture the visitor's video and audio. The video shows a middle-aged man, and the audio is mixed with unnatural noise. The captured data is immediately sent to the server using a secure protocol, where it is analyzed. The analysis results indicate that the visitor is estimated to be 45 years old and that he has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background. Based on this information, the server calculates a high score and determines the visitor to be suspicious. Finally, a notification method is activated, sending an alert to the user's smartphone and also notifying the police.
[0486] Prompt Sentence Examples
[0487] "The video of a middle-aged man captured by an intercom camera contains unnatural noise. We will analyze this visitor to determine whether he is suspicious."
[0488] This system allows users to monitor visitor information in real time to ensure safety, and also provides effective crime prevention measures for vulnerable groups such as the elderly.
[0489] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0490] Step 1:
[0491] The device uses the intercom's camera and microphone to capture video and audio of visitors in real time. The camera automatically captures video when the visitor stands in front of the intercom, and the microphone records audio when the visitor speaks. The captured video and audio data are then time-stamped.
[0492] Input: Visitor video and audio
[0493] Output: Time-stamped video and audio data
[0494] How it works: When a visitor presses the button on the intercom, the camera and microphone automatically activate and record video and audio.
[0495] Step 2:
[0496] The device transmits the captured, time-stamped video and audio data to the server using a secure protocol (e.g., SSL / TLS), where the data is buffered and transmitted in real time.
[0497] Input: Time-stamped video and audio data
[0498] Output: Securely transmitted video and audio data
[0499] What happens: The device buffers the captured data and prepares to send it to the server over a secure channel. The transmission begins immediately.
[0500] Step 3:
[0501] The server uses a data analysis module to analyze the video and audio data it receives. First, video analysis involves facial recognition and age and gender estimation, while audio analysis involves the presence or absence of background noise and the characteristics of the audio.
[0502] Input: Securely transmitted video and audio data
[0503] Output: Visitor demographic data (e.g., age, gender, presence or absence of background noise)
[0504] How it works: The server analyzes the video data frame by frame to extract facial features, and simultaneously performs spectrogram analysis of the audio data to check for abnormal noise.
[0505] Step 4:
[0506] The server uses a score calculation module to calculate a score based on attributes extracted from the analysis results (e.g., the visitor's age, gender, whether they are a known person, whether there is background noise, etc.). The score is weighted for each attribute, and the overall likelihood of a suspicious person is evaluated.
[0507] Input: Visitor attribute data
[0508] Output: The calculated score
[0509] What happens: The server runs a program that estimates the visitor's age to be 45, determines their gender as male, checks for the presence of background noise, and calculates a score based on these attributes.
[0510] Step 5:
[0511] The server determines whether the visitor is suspicious based on the calculated score. If the visitor is determined to be suspicious, the server proceeds to the next notification process.
[0512] Input: Calculated score
[0513] Output: Result of suspicious person judgment (e.g., suspicious person / not suspicious person)
[0514] Specific operation: If the score exceeds a certain threshold, the server executes logic to determine the person as suspicious.
[0515] Step 6:
[0516] The notification mechanism will be activated, sending an alert to the user's smartphone and, if necessary, notifying the police, allowing the user to quickly understand the situation and take action.
[0517] Input: Suspicious person detection result
[0518] Output: Notification to user and police
[0519] Specific operation: The server sends a warning to the user's smartphone via SMS or app notification that there is a "suspicious person" and also notifies the police using an API.
[0520] (Application example 1)
[0521] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0522] Conventional visitor reception systems are good at collecting visitor information, but lack the ability to effectively analyze the collected information and quickly determine whether a visitor is suspicious. Furthermore, they lack the means to issue real-time warnings or reports, making it difficult to prevent crimes or take early action. The present invention aims to solve these problems by providing a system that prevents crimes committed by visitors and ensures the safety of users.
[0523] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0524] In this invention, the server includes means for capturing video of visitors with a camera, means for capturing audio of visitors with a microphone, analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes, means for evaluating the visitor attributes based on the extracted attributes and calculating a score, means for determining whether the visitor is a suspicious person based on the calculated score, means for notifying the user and sending a warning to the police if the visitor is determined to be suspicious, means for linking the intercom and smart device to transfer visitor information in real time, and means for analyzing visitor attributes in real time and calculating a score. This allows for quick and accurate analysis of visitor information, enabling the detection of suspicious people and immediate response.
[0525] A "camera" is a device that captures video of visitors.
[0526] A "microphone" is a device that captures the visitor's voice.
[0527] A "generative artificial intelligence model" is an analytical method that analyzes captured video and audio to extract visitor attributes.
[0528] The "score calculation means" is a means having a function of evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[0529] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[0530] "Notification means" refers to the means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[0531] "Means for linking an intercom with a smart device" refers to a means for linking an intercom with a smart device such as a smartphone to transfer visitor information in real time.
[0532] The "real-time analysis means" is a means for analyzing visitor attributes in real time and calculating scores.
[0533] The system for implementing the present invention captures video and audio of visitors in real time and analyzes them using a generative artificial intelligence model to detect suspicious individuals and issue immediate warnings. This system uses the following hardware and software:
[0534] Hardware and Software Configuration
[0535] Hardware
[0536] Camera: A device that captures video of visitors, usually built into the intercom.
[0537] Microphone: A device that captures the visitor's voice and is built into the intercom, just like a camera.
[0538] Smart device: Usually a smartphone, which works in conjunction with the intercom to notify the user of visitor information.
[0539] software
[0540] Generative artificial intelligence model: An AI model that analyzes captured video and audio to extract visitor attributes (e.g., age, gender, background noise), for example using a deep learning framework such as Keras.
[0541] Score calculation algorithm: An algorithm for calculating the score based on the extracted attributes. It is implemented using a programming language such as Python.
[0542] Notification systems: Systems that send alerts to users when a visitor is deemed suspicious. This includes email sending services (such as smtplib) and mobile notification systems.
[0543] Process Overview
[0544] 1. The device uses a camera and microphone to capture video and audio of visitors in real time, and the captured data is immediately transmitted to an analytical tool that includes a generative artificial intelligence model.
[0545] 2. The server analyzes the video and audio data using a generative artificial intelligence model, which extracts attributes such as the visitor's age, gender, and the presence or absence of background noise.
[0546] 3. The server uses a scoring algorithm to calculate the visitor's score based on the extracted attributes, specifically, unknown visitors and those with unnatural noise are given a higher score.
[0547] 4. Based on the calculated score, the server determines whether the visitor is suspicious.
[0548] 5. If a score exceeds the threshold specified by the user, the system will identify the person as suspicious and immediately send a warning to the user's smart device. If requested, the system will also automatically notify the police.
[0549] Specific examples
[0550] For example, if a middle-aged man is seen on camera as a visitor and unnatural noise is detected in the background, the generative AI model will estimate his age to be 45 and detect background noise from the audio data. Based on this information, the scoring means will assign a high score and determine that the visitor is suspicious. The user will then receive a real-time notification saying, "A suspicious visitor has been detected!"
[0551] An example of an input prompt for a generative AI model is as follows:
[0552] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[0553] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0554] Step 1:
[0555] The device uses a camera and microphone to capture video and audio of visitors in real time.
[0556] Input: Real-time video data, real-time audio data
[0557] Data processing: Video and audio data is acquired and converted into digital format.
[0558] Output: Digital video data, digital audio data
[0559] Step 2:
[0560] The terminal transfers the captured video and audio data to the server.
[0561] Input: Digital video data, digital audio data
[0562] Data processing: Packetizing data and transferring it over the network
[0563] Output: Video and audio data transferred to the server
[0564] Step 3:
[0565] The server uses a generative artificial intelligence model to analyze the captured video and audio data.
[0566] Input: Video data, audio data
[0567] Data processing: Data analysis using generative AI models (extracting attributes such as age, gender, and background noise)
[0568] Output: Visitor demographic data (age, gender, presence of background noise, etc.)
[0569] Step 4:
[0570] The server uses a scoring algorithm to calculate a score for the visitor based on the extracted attributes.
[0571] Input: Visitor attribute data
[0572] Data processing: Score calculation algorithm (adds score according to attributes)
[0573] Output: Visitor score
[0574] Step 5:
[0575] The server determines whether the visitor is suspicious based on the calculated score.
[0576] Input: Visitor score
[0577] Data processing: Comparing scores with set thresholds
[0578] Output: Suspicious person determination result (suspicious person or non-suspicious person)
[0579] Step 6:
[0580] If a person is determined to be suspicious, the server will send a warning to the user's smart device and, if necessary, notify the police.
[0581] Input: Suspicious person judgment result
[0582] Data processing: generating and sending notification messages (contacting users and the police)
[0583] Output: Warning notice to user, report to police
[0584] Considering the following specific example, an example prompt would be:
[0585] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[0586] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0587] This invention is a system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[0588] System Configuration
[0589] The present invention comprises the following main components:
[0590] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0591] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0592] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0593] 4. The emotion engine is a means of recognizing the user's emotions, and uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions.
[0594] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[0595] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0596] 7. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[0597] System Operation
[0598] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[0599] As a specific example, the following scenario can be considered.
[0600] Example scenario:
[0601] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0602] 2. The video shows a middle-aged woman, and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice. In front of the intercom, the user's face takes on an anxious expression and her voice becomes trembling.
[0603] 3. The server analyzes the data, estimates the visitor's age to be 35, and confirms that the person has not previously been registered as an acquaintance. Furthermore, audio analysis detects no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[0604] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[0605] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[0606] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0607] In this way, users can be more effectively protected from specialized fraud and other crimes by using a crime prevention system that combines emotion analysis, and the system can provide more advanced security measures to a wider range of people, including the elderly.
[0608] The processing flow will be explained below.
[0609] Step 1:
[0610] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[0611] Step 2:
[0612] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[0613] Step 3:
[0614] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[0615] Step 4:
[0616] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0617] Step 5:
[0618] The device also captures the user's facial expressions and tone of voice to recognize their emotions. An emotion engine is used to analyze the user's emotions and identify emotions such as anxiety or fear.
[0619] Step 6:
[0620] The server calculates a score based on the extracted visitor attributes and the user's perceived emotions. For example, if an elderly person or unnatural noise is detected, the score will be higher. The score will also increase if the user shows signs of anxiety or fear.
[0621] Step 7:
[0622] The server determines whether the visitor is suspicious based on the calculated total score. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[0623] Step 8:
[0624] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[0625] Step 9:
[0626] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[0627] Step 10:
[0628] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[0629] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[0630] Example 2
[0631] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0632] There is a need to quickly and accurately determine whether a visitor is suspicious and ensure the safety of users. Conventional systems only analyze visitor attributes and are unable to integrate information including user emotions. As a result, there is a problem that the accuracy of detecting suspicious individuals is low and users' anxiety cannot be completely alleviated.
[0633] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0634] In this invention, the server includes a means for capturing video of the visitor with a camera, a means for capturing the visitor's voice with a microphone, and a means for capturing the user's facial expression and tone of voice and analyzing the user's emotions, thereby making it possible to determine suspicious individuals by taking into account not only the visitor's attributes but also the user's emotions.
[0635] A "camera" is a device for capturing images.
[0636] A "microphone" is a device for capturing sound.
[0637] A "generative artificial intelligence model" is an artificial intelligence technology that analyzes specific features and patterns from input data and outputs the results.
[0638] "Analysis means" refers to a device or software for analyzing the captured video and audio data and extracting visitor attributes.
[0639] "Means for capturing a user's facial expression and tone of voice and analyzing the user's emotions" refers to a device or software that uses a camera and microphone to capture a user's facial expression and tone of voice and, based on that, identifies the user's emotions.
[0640] The "means for calculating a score" is a device or software for evaluating visitor attributes based on the extracted visitor attributes and the analyzed user sentiment.
[0641] The "means for determining" is a device or software for determining whether a visitor is suspicious or not based on the calculated score.
[0642] The "notification means" is a device or system that notifies the user when a person is determined to be suspicious and also sends a warning to the police.
[0643] This invention is a security system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[0644] System Configuration
[0645] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0646] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0647] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data, which extracts attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. The generative artificial intelligence model used is a commonly used AI model (such as OpenAI's GPT-4).
[0648] 4. The emotion engine is a means of recognizing the user's emotions, and analyzes the user's facial expressions and tone of voice using a camera and microphone. Emotion engines that can be used include Amazon's AWS Rekognition and Google Cloud Vision API.
[0649] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[0650] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0651] 7. If a person is determined to be suspicious, the notification method will notify the user and also send an alert to the police. This method allows the user to take appropriate action quickly. Services such as Twilio and Pushbullet can be used as notification methods.
[0652] System Operation
[0653] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[0654] Specific examples
[0655] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0656] 2. The video shows a middle-aged man and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice.
[0657] 3. The server analyzes the data, estimates the visitor's age to be 45, and confirms that this person has not previously been registered as an acquaintance. Furthermore, audio analysis detects that there is no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[0658] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[0659] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[0660] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0661] Examples of prompts:
[0662] Analyze the following data:
[0663] Visitor video data, age, gender, and whether they are known faces
[0664] Visitor voice data and whether there are any abnormal sounds in the background
[0665] Analysis results of user's facial expressions and tone of voice
[0666] The analysis methods are as follows:
[0667] 1. Identifying visitor attributes based on video data
[0668] 2. Analyze background sounds based on audio data
[0669] 3. Analyze user emotions
[0670] Calculate your overall score based on the analysis results.
[0671] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0672] Step 1:
[0673] Data Capture
[0674] The device captures the visitor's video and audio in real time. Specifically, the intercom captures video with a built-in camera and audio with a built-in microphone. The input is the visitor's real-time video and audio, and is used to output video and audio data.
[0675] Specific behavior:
[0676] The device's camera will then begin working, capturing high-resolution footage of the visitor.
[0677] The device's microphone will begin working and will record the visitor's speech and background sounds.
[0678] The captured video and audio data is temporarily stored in the device's memory.
[0679] Step 2:
[0680] Data transmission
[0681] The device sends the captured video and audio data to the server. The input is the captured video and audio data, which is then encrypted and sent to the server for output.
[0682] Specific behavior:
[0683] The terminal encrypts the video and audio data and transmits it to a server via a communications network.
[0684] Check whether the data was sent successfully and proceed to the next step.
[0685] Step 3:
[0686] Visitor attribute analysis
[0687] The server analyzes the received video data using a generative artificial intelligence model. The input is the video and audio data received from the device, and the output is the visitor's attribute data, such as age, gender, and whether or not the face is familiar.
[0688] Specific behavior:
[0689] The server inputs video data into the generative artificial intelligence model using prompt sentences and acquires attributes such as age and gender.
[0690] The server analyzes the audio data and determines whether there are any suspicious noises in the background sound.
[0691] The analysis results are stored in a database on the server.
[0692] Step 4:
[0693] User sentiment analysis
[0694] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and sends them to the server. The input is the user's facial expression data and tone of voice data, which is then used by the emotion engine to analyze the user's emotions and output them.
[0695] Specific behavior:
[0696] The device's camera captures the user's facial expressions in high resolution.
[0697] The device's microphone records the user's voice.
[0698] The captured data is sent to a server and analyzed by an emotion engine.
[0699] The analysis results, which represent the user's emotions (e.g., anxiety, fear, relief), are output to the server.
[0700] Step 5:
[0701] Calculation of overall score
[0702] The server calculates an overall score using the score calculation means based on the data obtained from the analysis means and emotion engine. The input is visitor attribute data and user emotion data, which are integrated to calculate and output an overall score.
[0703] Specific behavior:
[0704] The server combines visitor attribute data (e.g., age, gender, whether the face is familiar) with user emotion data (e.g., anxiety).
[0705] A score calculation means is used to quantify the overall score.
[0706] The calculated overall score is stored in a database on the server.
[0707] Step 6:
[0708] judgement
[0709] The server determines whether the visitor is suspicious or not based on the calculated total score. The input is the total score, and the output is the result of the determination whether the visitor is suspicious or not.
[0710] Specific behavior:
[0711] The server compares the overall score to a threshold.
[0712] If the score exceeds the threshold, the visitor is deemed suspicious.
[0713] The judgment results are stored in the server's database.
[0714] Step 7:
[0715] Notification and Response
[0716] If the visitor is determined to be suspicious, the device automatically responds to the visitor saying "We are unable to serve you," and the server sends a warning to the user's smartphone and even notifies the police. The input is the judgment result, and the output is the execution of the notification and report.
[0717] Specific behavior:
[0718] The device will send a voice notification to the visitor via the intercom saying "We are unavailable."
[0719] The server uses a service such as Twilio or Pushbullet to send an alert message to the user's smartphone.
[0720] The server notifies the police using a preset notification means.
[0721] (Application example 2)
[0722] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0723] In modern society, special frauds and other criminal activities are on the rise, posing a serious problem especially for the elderly and people living alone. Therefore, there is a need for advanced security systems that can efficiently analyze visitor identities and user emotions in real time to prevent criminal activities. However, conventional systems lack a comprehensive judgment function that takes into account not only visitor attributes but also user emotions, resulting in low fraud detection accuracy.
[0724] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0725] In this invention, the server includes: means for capturing video of visitors with a camera; means for capturing audio of visitors with a microphone; analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes; means for analyzing the user's video and audio to recognize the user's emotions; means for evaluating the visitor's attributes based on the extracted attributes and the recognized user emotions and calculating a score; means for determining whether the visitor is a suspicious person based on the calculated score; and means for notifying the user and sending a warning to the police if the visitor is determined to be a suspicious person. This realizes a comprehensive security system that integrates visitor identity analysis and user emotion recognition, making it possible to prevent specialized frauds and other criminal acts.
[0726] A "camera" is a device that captures video of visitors.
[0727] A "microphone" is a device that captures the visitor's voice.
[0728] A "generative artificial intelligence model" is a type of artificial intelligence that analyzes captured video and audio to extract visitor attributes.
[0729] "Analysis means" refers to means for analyzing captured video and audio, including generative artificial intelligence models.
[0730] A "user" is a person who uses this system.
[0731] The "means for recognizing emotions" is a means for identifying the emotions of a user by analyzing the user's video and audio.
[0732] The "means for evaluating attributes and calculating a score" is a means for evaluating attributes of a visitor based on the extracted attributes and the recognized user sentiment, and calculating a score.
[0733] The "means for determining whether a visitor is suspicious" is a means for determining whether a visitor is suspicious based on the calculated score.
[0734] The "means for notifying and sending a warning" is a means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[0735] This invention presents a form of advanced security system that analyzes visitor identity and user sentiment. The system consists of the following main components:
[0736] Hardware:
[0737] The system's terminal is a device equipped with a camera and microphone. This device functions as an intercom and captures real-time video and audio information of visitors, allowing the terminal to accurately record their actions and comments.
[0738] software:
[0739] The system applies a generative artificial intelligence model (AI model) and an emotion recognition engine that analyzes captured video and audio data to extract visitor attributes (age, gender, known or unknown) and user emotions (anxiety, relief, fear, etc.).
[0740] Analysis method:
[0741] Using AI models, the video and audio captured on the device are analyzed to extract visitor attributes.
[0742] Emotion recognition means:
[0743] The camera and microphone are used to analyze the user's facial expressions and tone of voice, and the emotion recognition engine identifies the user's emotions.
[0744] Data processing and calculation:
[0745] The server calculates a total score using the score calculation means based on the visitor's attributes extracted by the analysis means and the user's emotions identified by the emotion recognition means. This total score is used to determine whether the visitor is suspicious. Once a determination is made, the notification means is activated and a warning is sent to the user and the police.
[0746] Examples:
[0747] When a visitor standing at the front door operates the intercom, the camera automatically captures video and the microphone captures audio. This data is analyzed by an AI model to determine the visitor's age, gender, and whether they are a familiar face. Meanwhile, the user's camera video and audio are analyzed by an emotion recognition engine to identify the user's emotions. Based on this, the server calculates an overall score, integrating the visitor's attributes and the user's emotions to determine whether they are suspicious.
[0748] Example prompt sentence:
[0749] "Capture video of visitors and analyze their attributes."
[0750] "Estimate the visitor's age and gender and assess user sentiment."
[0751] "Use a comprehensive score to determine whether a visitor is suspicious."
[0752] In this way, the system provides comprehensive security that integrates visitor identity analysis and user emotion recognition, making it possible to effectively protect users from specialized fraud and other criminal activities.
[0753] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0754] Step 1:
[0755] The device uses a camera and microphone to capture the visitor's video and audio. The input is the visitor's real-time video and audio, and the output is the captured digital data. This digital data is temporarily stored for subsequent analysis processes.
[0756] Step 2:
[0757] The terminal sends the captured video and audio data to the analysis means. The input is the captured video and audio digital data, and the output is the visitor's attribute data (e.g., age, gender, whether the visitor is a known person) as the analysis result. The terminal analyzes the data using a generative artificial intelligence model and extracts the attributes.
[0758] Step 3:
[0759] The terminal captures the user's video and audio. This step is performed between the time the visitor operates the intercom and the time the user answers. The input is the user's real-time video and audio, and the output is the captured digital data.
[0760] Step 4:
[0761] The device sends the captured user video and audio data to the emotion recognition means. The input is the captured user video and audio digital data, and the output is the user's emotional data (e.g., anxiety, relief, fear, etc.) as an analysis result. The emotion recognition engine analyzes this data to identify the user's emotions.
[0762] Step 5:
[0763] The server integrates the output data of the analysis means and emotion recognition means and calculates an overall score using the score calculation means. The input is the visitor's attribute data and the user's emotion data, and the output is an overall score. The server performs data calculations based on this data and evaluates whether the visitor is suspicious or not.
[0764] Step 6:
[0765] The server determines whether the visitor is suspicious based on the total score. The input is the total score, and the output is the result of the judgment whether the visitor is suspicious. Based on this judgment result, a notification method is prepared.
[0766] Step 7:
[0767] If the server determines that a person is suspicious, it immediately notifies the user and sends a warning to the police. The input is the suspicious person determination result, and the output is a notification message to the user and a warning message to the police. The server automatically generates and sends the message.
[0768] In this way, the system analyzes the identity of visitors and the user's emotions in real time, enabling it to identify suspicious individuals and take prompt action.
[0769] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0770] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0771] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0772] [Third embodiment]
[0773] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0774] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0775] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0776] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0777] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0778] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0779] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0780] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0781] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0782] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0783] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0784] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0785] This invention is a system for preventing special fraud and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and aims to detect visitors who may be committing fraud.
[0786] System Configuration
[0787] The present invention comprises the following main components:
[0788] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0789] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0790] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0791] 4. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates a score. The score is higher for elderly people, unfamiliar faces, and unnatural background sounds.
[0792] 5. The judgment means judges whether the visitor is suspicious or not based on the calculated score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0793] 6. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[0794] System Operation
[0795] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. The server then evaluates the extracted attributes using a score calculation means to calculate a score. Based on this score, the server determines whether the visitor is suspicious.
[0796] As a specific example, the following scenario can be considered.
[0797] Example scenario:
[0798] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0799] 2. The video shows a middle-aged man and the audio is mixed with unnatural noise. The video and audio data are automatically sent to an analysis device.
[0800] 3. The server analyzes the data, estimates the visitor's age to be 45, and verifies that the person has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background.
[0801] 4. Based on these attributes, the server calculates a score. For example, a high score can be set if the person's age is within a certain range, if they are not registered as an acquaintance, or if unnatural noise is detected.
[0802] 5. The visitor is determined to be suspicious based on the score calculated by the server.
[0803] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0804] By using such a system, users can protect themselves from crime, and it can provide an effective crime prevention measure, especially for those who are more likely to become victims of crime, such as the elderly.
[0805] The processing flow will be explained below.
[0806] Step 1:
[0807] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[0808] Step 2:
[0809] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[0810] Step 3:
[0811] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[0812] Step 4:
[0813] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0814] Step 5:
[0815] The server calculates a score for each visitor based on the extracted attributes. For example, an elderly person will receive a higher score, a familiar face will receive a lower score, and the presence of unnatural background noise will also result in a higher score.
[0816] Step 6:
[0817] The server uses the calculated score to determine whether the visitor is suspicious. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[0818] Step 7:
[0819] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[0820] Step 8:
[0821] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[0822] Step 9:
[0823] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[0824] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[0825] Example 1
[0826] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0827] In recent years, there has been an increase in special frauds and other crimes, with the elderly and other socially vulnerable people being particularly vulnerable. Conventional crime prevention measures monitor visitor information in real time, but it is often difficult to accurately identify suspicious individuals. Furthermore, ensuring the security of captured data and linking it to a prompt notification system for suspicious individuals have been issues.
[0828] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0829] In this invention, the server includes a means for capturing video of visitors using a camera, a means for capturing audio of visitors using a microphone, and a means for transmitting the captured and analyzed data using a secure protocol. This allows visitor information to be captured in real time, securely transmitted, and analyzed, making it possible to accurately evaluate visitor attributes and identify suspicious individuals. This also ensures data security and allows for prompt notification of suspicious individuals.
[0830] A "camera" is a device that captures video of visitors.
[0831] A "microphone" is a device that captures the visitor's voice.
[0832] A "generative artificial intelligence model" is an artificial intelligence model that analyzes captured video and audio data and extracts visitor attributes.
[0833] The "analysis means" is a means for analyzing the captured video and audio data to extract visitor attributes.
[0834] "Attributes" refers to characteristics such as the visitor's age, gender, whether they are a known person, and any abnormalities in background noise.
[0835] The "score calculation means" is a means for evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[0836] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[0837] The "notification means" is a means for notifying the user and the police if a person is determined to be a suspicious person.
[0838] A "secure protocol" is a communication method that ensures security in data transmission.
[0839] A "timestamp" is time information added to captured data.
[0840] A "data analysis module" is a piece of software functionality for extracting attributes from captured video and audio data.
[0841] A "score calculation module" is a part of the software function that calculates a score based on visitor attributes.
[0842] A "judgment module" is a part of the software function that judges whether a visitor is suspicious based on the calculated score.
[0843] This invention is a system for preventing special fraud and other crimes committed by visitors. The system is configured using the following hardware and software:
[0844] Hardware
[0845] 1. Camera: This is a device built into the intercom that captures video of visitors. When a visitor stands in front of the intercom, the camera automatically starts recording.
[0846] 2. Microphone: This is a device built into the intercom that captures the visitor's voice. When a visitor speaks, the microphone records the voice.
[0847] software
[0848] 1. Generative AI model: An AI model that analyzes captured video and audio data and extracts visitor attributes.
[0849] 2. Analysis method: A method for analyzing captured video and audio data to extract visitor attributes.
[0850] 3. Score calculation means: A means for evaluating the visitor's attributes based on the extracted attributes and calculating a score.
[0851] 4. Judgment method: A method for judging whether a visitor is suspicious or not based on the calculated score.
[0852] 5. Notification means: A means for notifying the user and the police based on the judgment results.
[0853] 6. Secure Protocol: A communication method that ensures security in data transmission.
[0854] 7. Data Analysis Module: This is the part of the software that functions to extract attributes from the captured video and audio data.
[0855] 8. Score calculation module: A part of the software function that calculates a score based on visitor attributes.
[0856] 9. Judgment module: This is the part of the software function that judges whether a visitor is suspicious or not based on the calculated score.
[0857] The system works as follows: First, the terminal captures the visitor's video and audio in real time using the intercom's built-in camera and microphone. Next, the captured data is sent to the server using a secure protocol (e.g., SSL / TLS). The server then analyzes the received video and audio data using a generative artificial intelligence model to extract attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. These attributes are evaluated using a score calculation means to calculate a score.
[0858] As a specific example, when a visitor presses the button on the intercom, the device activates the camera and microphone to capture the visitor's video and audio. The video shows a middle-aged man, and the audio is mixed with unnatural noise. The captured data is immediately sent to the server using a secure protocol, where it is analyzed. The analysis results indicate that the visitor is estimated to be 45 years old and that he has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background. Based on this information, the server calculates a high score and determines the visitor to be suspicious. Finally, a notification method is activated, sending an alert to the user's smartphone and also notifying the police.
[0859] Prompt Sentence Examples
[0860] "The video of a middle-aged man captured by an intercom camera contains unnatural noise. We will analyze this visitor to determine whether he is suspicious."
[0861] This system allows users to monitor visitor information in real time to ensure safety, and also provides effective crime prevention measures for vulnerable groups such as the elderly.
[0862] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0863] Step 1:
[0864] The device uses the intercom's camera and microphone to capture video and audio of visitors in real time. The camera automatically captures video when the visitor stands in front of the intercom, and the microphone records audio when the visitor speaks. The captured video and audio data are then time-stamped.
[0865] Input: Visitor video and audio
[0866] Output: Time-stamped video and audio data
[0867] How it works: When a visitor presses the button on the intercom, the camera and microphone automatically activate and record video and audio.
[0868] Step 2:
[0869] The device transmits the captured, time-stamped video and audio data to the server using a secure protocol (e.g., SSL / TLS), where the data is buffered and transmitted in real time.
[0870] Input: Time-stamped video and audio data
[0871] Output: Securely transmitted video and audio data
[0872] What happens: The device buffers the captured data and prepares to send it to the server over a secure channel. The transmission begins immediately.
[0873] Step 3:
[0874] The server uses a data analysis module to analyze the video and audio data it receives. First, video analysis involves facial recognition and age and gender estimation, while audio analysis involves the presence or absence of background noise and the characteristics of the audio.
[0875] Input: Securely transmitted video and audio data
[0876] Output: Visitor demographic data (e.g., age, gender, presence or absence of background noise)
[0877] How it works: The server analyzes the video data frame by frame to extract facial features, and simultaneously performs spectrogram analysis of the audio data to check for abnormal noise.
[0878] Step 4:
[0879] The server uses a score calculation module to calculate a score based on attributes extracted from the analysis results (e.g., the visitor's age, gender, whether they are a known person, whether there is background noise, etc.). The score is weighted for each attribute, and the overall likelihood of a suspicious person is evaluated.
[0880] Input: Visitor attribute data
[0881] Output: The calculated score
[0882] What happens: The server runs a program that estimates the visitor's age to be 45, determines their gender as male, checks for the presence of background noise, and calculates a score based on these attributes.
[0883] Step 5:
[0884] The server determines whether the visitor is suspicious based on the calculated score. If the visitor is determined to be suspicious, the server proceeds to the next notification process.
[0885] Input: Calculated score
[0886] Output: Result of suspicious person judgment (e.g., suspicious person / not suspicious person)
[0887] Specific operation: If the score exceeds a certain threshold, the server executes logic to determine the person as suspicious.
[0888] Step 6:
[0889] The notification mechanism will be activated, sending an alert to the user's smartphone and, if necessary, notifying the police, allowing the user to quickly understand the situation and take action.
[0890] Input: Suspicious person detection result
[0891] Output: Notification to user and police
[0892] Specific operation: The server sends a warning to the user's smartphone via SMS or app notification that there is a "suspicious person" and also notifies the police using an API.
[0893] (Application example 1)
[0894] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0895] Conventional visitor reception systems are good at collecting visitor information, but lack the ability to effectively analyze the collected information and quickly determine whether a visitor is suspicious. Furthermore, they lack the means to issue real-time warnings or reports, making it difficult to prevent crimes or take early action. The present invention aims to solve these problems by providing a system that prevents crimes committed by visitors and ensures the safety of users.
[0896] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0897] In this invention, the server includes means for capturing video of visitors with a camera, means for capturing audio of visitors with a microphone, analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes, means for evaluating the visitor attributes based on the extracted attributes and calculating a score, means for determining whether the visitor is a suspicious person based on the calculated score, means for notifying the user and sending a warning to the police if the visitor is determined to be suspicious, means for linking the intercom and smart device to transfer visitor information in real time, and means for analyzing visitor attributes in real time and calculating a score. This allows for quick and accurate analysis of visitor information, enabling the detection of suspicious people and immediate response.
[0898] A "camera" is a device that captures video of visitors.
[0899] A "microphone" is a device that captures the visitor's voice.
[0900] A "generative artificial intelligence model" is an analytical method that analyzes captured video and audio to extract visitor attributes.
[0901] The "score calculation means" is a means having a function of evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[0902] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[0903] "Notification means" refers to the means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[0904] "Means for linking an intercom with a smart device" refers to a means for linking an intercom with a smart device such as a smartphone to transfer visitor information in real time.
[0905] The "real-time analysis means" is a means for analyzing visitor attributes in real time and calculating scores.
[0906] The system for implementing the present invention captures video and audio of visitors in real time and analyzes them using a generative artificial intelligence model to detect suspicious individuals and issue immediate warnings. This system uses the following hardware and software:
[0907] Hardware and Software Configuration
[0908] Hardware
[0909] Camera: A device that captures video of visitors, usually built into the intercom.
[0910] Microphone: A device that captures the visitor's voice and is built into the intercom, just like a camera.
[0911] Smart device: Usually a smartphone, which works in conjunction with the intercom to notify the user of visitor information.
[0912] software
[0913] Generative artificial intelligence model: An AI model that analyzes captured video and audio to extract visitor attributes (e.g., age, gender, background noise), for example using a deep learning framework such as Keras.
[0914] Score calculation algorithm: An algorithm for calculating the score based on the extracted attributes. It is implemented using a programming language such as Python.
[0915] Notification systems: Systems that send alerts to users when a visitor is deemed suspicious. This includes email sending services (such as smtplib) and mobile notification systems.
[0916] Process Overview
[0917] 1. The device uses a camera and microphone to capture video and audio of visitors in real time, and the captured data is immediately transmitted to an analytical tool that includes a generative artificial intelligence model.
[0918] 2. The server analyzes the video and audio data using a generative artificial intelligence model, which extracts attributes such as the visitor's age, gender, and the presence or absence of background noise.
[0919] 3. The server uses a scoring algorithm to calculate the visitor's score based on the extracted attributes, specifically, unknown visitors and those with unnatural noise are given a higher score.
[0920] 4. Based on the calculated score, the server determines whether the visitor is suspicious.
[0921] 5. If a score exceeds the threshold specified by the user, the system will identify the person as suspicious and immediately send a warning to the user's smart device. If requested, the system will also automatically notify the police.
[0922] Specific examples
[0923] For example, if a middle-aged man is seen on camera as a visitor and unnatural noise is detected in the background, the generative AI model will estimate his age to be 45 and detect background noise from the audio data. Based on this information, the scoring means will assign a high score and determine that the visitor is suspicious. The user will then receive a real-time notification saying, "A suspicious visitor has been detected!"
[0924] An example of an input prompt for a generative AI model is as follows:
[0925] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[0926] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0927] Step 1:
[0928] The device uses a camera and microphone to capture video and audio of visitors in real time.
[0929] Input: Real-time video data, real-time audio data
[0930] Data processing: Video and audio data is acquired and converted into digital format.
[0931] Output: Digital video data, digital audio data
[0932] Step 2:
[0933] The terminal transfers the captured video and audio data to the server.
[0934] Input: Digital video data, digital audio data
[0935] Data processing: Packetizing data and transferring it over the network
[0936] Output: Video and audio data transferred to the server
[0937] Step 3:
[0938] The server uses a generative artificial intelligence model to analyze the captured video and audio data.
[0939] Input: Video data, audio data
[0940] Data processing: Data analysis using generative AI models (extracting attributes such as age, gender, and background noise)
[0941] Output: Visitor demographic data (age, gender, presence of background noise, etc.)
[0942] Step 4:
[0943] The server uses a scoring algorithm to calculate a score for the visitor based on the extracted attributes.
[0944] Input: Visitor attribute data
[0945] Data processing: Score calculation algorithm (adds score according to attributes)
[0946] Output: Visitor score
[0947] Step 5:
[0948] The server determines whether the visitor is suspicious based on the calculated score.
[0949] Input: Visitor score
[0950] Data processing: Comparing scores with set thresholds
[0951] Output: Suspicious person determination result (suspicious person or non-suspicious person)
[0952] Step 6:
[0953] If a person is determined to be suspicious, the server will send a warning to the user's smart device and, if necessary, notify the police.
[0954] Input: Suspicious person judgment result
[0955] Data processing: generating and sending notification messages (contacting users and the police)
[0956] Output: Warning notice to user, report to police
[0957] Considering the following specific example, an example prompt would be:
[0958] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[0959] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0960] This invention is a system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[0961] System Configuration
[0962] The present invention comprises the following main components:
[0963] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[0964] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[0965] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0966] 4. The emotion engine is a means of recognizing the user's emotions, and uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions.
[0967] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[0968] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[0969] 7. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[0970] System Operation
[0971] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[0972] As a specific example, the following scenario can be considered.
[0973] Example scenario:
[0974] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[0975] 2. The video shows a middle-aged woman, and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice. In front of the intercom, the user's face takes on an anxious expression and her voice becomes trembling.
[0976] 3. The server analyzes the data, estimates the visitor's age to be 35, and confirms that the person has not previously been registered as an acquaintance. Furthermore, audio analysis detects no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[0977] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[0978] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[0979] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[0980] In this way, users can be more effectively protected from specialized fraud and other crimes by using a crime prevention system that combines emotion analysis, and the system can provide more advanced security measures to a wider range of people, including the elderly.
[0981] The processing flow will be explained below.
[0982] Step 1:
[0983] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[0984] Step 2:
[0985] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[0986] Step 3:
[0987] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[0988] Step 4:
[0989] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[0990] Step 5:
[0991] The device also captures the user's facial expressions and tone of voice to recognize their emotions. An emotion engine is used to analyze the user's emotions and identify emotions such as anxiety or fear.
[0992] Step 6:
[0993] The server calculates a score based on the extracted visitor attributes and the user's perceived emotions. For example, if an elderly person or unnatural noise is detected, the score will be higher. The score will also increase if the user shows signs of anxiety or fear.
[0994] Step 7:
[0995] The server determines whether the visitor is suspicious based on the calculated total score. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[0996] Step 8:
[0997] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[0998] Step 9:
[0999] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[1000] Step 10:
[1001] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[1002] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[1003] Example 2
[1004] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1005] There is a need to quickly and accurately determine whether a visitor is suspicious and ensure the safety of users. Conventional systems only analyze visitor attributes and are unable to integrate information including user emotions. As a result, there is a problem that the accuracy of detecting suspicious individuals is low and users' anxiety cannot be completely alleviated.
[1006] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1007] In this invention, the server includes a means for capturing video of the visitor with a camera, a means for capturing the visitor's voice with a microphone, and a means for capturing the user's facial expression and tone of voice and analyzing the user's emotions, thereby making it possible to determine suspicious individuals by taking into account not only the visitor's attributes but also the user's emotions.
[1008] A "camera" is a device for capturing images.
[1009] A "microphone" is a device for capturing sound.
[1010] A "generative artificial intelligence model" is an artificial intelligence technology that analyzes specific features and patterns from input data and outputs the results.
[1011] "Analysis means" refers to a device or software for analyzing the captured video and audio data and extracting visitor attributes.
[1012] "Means for capturing a user's facial expression and tone of voice and analyzing the user's emotions" refers to a device or software that uses a camera and microphone to capture a user's facial expression and tone of voice and, based on that, identifies the user's emotions.
[1013] The "means for calculating a score" is a device or software for evaluating visitor attributes based on the extracted visitor attributes and the analyzed user sentiment.
[1014] The "means for determining" is a device or software for determining whether a visitor is suspicious or not based on the calculated score.
[1015] The "notification means" is a device or system that notifies the user when a person is determined to be suspicious and also sends a warning to the police.
[1016] This invention is a security system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[1017] System Configuration
[1018] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[1019] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[1020] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data, which extracts attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. The generative artificial intelligence model used is a commonly used AI model (such as OpenAI's GPT-4).
[1021] 4. The emotion engine is a means of recognizing the user's emotions, and analyzes the user's facial expressions and tone of voice using a camera and microphone. Emotion engines that can be used include Amazon's AWS Rekognition and Google Cloud Vision API.
[1022] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[1023] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[1024] 7. If a person is determined to be suspicious, the notification method will notify the user and also send an alert to the police. This method allows the user to take appropriate action quickly. Services such as Twilio and Pushbullet can be used as notification methods.
[1025] System Operation
[1026] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[1027] Specific examples
[1028] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[1029] 2. The video shows a middle-aged man and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice.
[1030] 3. The server analyzes the data, estimates the visitor's age to be 45, and confirms that this person has not previously been registered as an acquaintance. Furthermore, audio analysis detects that there is no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[1031] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[1032] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[1033] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[1034] Examples of prompts:
[1035] Analyze the following data:
[1036] Visitor video data, age, gender, and whether they are known faces
[1037] Visitor voice data and whether there are any abnormal sounds in the background
[1038] Analysis results of user's facial expressions and tone of voice
[1039] The analysis methods are as follows:
[1040] 1. Identifying visitor attributes based on video data
[1041] 2. Analyze background sounds based on audio data
[1042] 3. Analyze user emotions
[1043] Calculate your overall score based on the analysis results.
[1044] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1045] Step 1:
[1046] Data Capture
[1047] The device captures the visitor's video and audio in real time. Specifically, the intercom captures video with a built-in camera and audio with a built-in microphone. The input is the visitor's real-time video and audio, and is used to output video and audio data.
[1048] Specific behavior:
[1049] The device's camera will then begin working, capturing high-resolution footage of the visitor.
[1050] The device's microphone will begin working and will record the visitor's speech and background sounds.
[1051] The captured video and audio data is temporarily stored in the device's memory.
[1052] Step 2:
[1053] Data transmission
[1054] The device sends the captured video and audio data to the server. The input is the captured video and audio data, which is then encrypted and sent to the server for output.
[1055] Specific behavior:
[1056] The terminal encrypts the video and audio data and transmits it to a server via a communications network.
[1057] Check whether the data was sent successfully and proceed to the next step.
[1058] Step 3:
[1059] Visitor attribute analysis
[1060] The server analyzes the received video data using a generative artificial intelligence model. The input is the video and audio data received from the device, and the output is the visitor's attribute data, such as age, gender, and whether or not the face is familiar.
[1061] Specific behavior:
[1062] The server inputs video data into the generative artificial intelligence model using prompt sentences and acquires attributes such as age and gender.
[1063] The server analyzes the audio data and determines whether there are any suspicious noises in the background sound.
[1064] The analysis results are stored in a database on the server.
[1065] Step 4:
[1066] User sentiment analysis
[1067] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and sends them to the server. The input is the user's facial expression data and tone of voice data, which is then used by the emotion engine to analyze the user's emotions and output them.
[1068] Specific behavior:
[1069] The device's camera captures the user's facial expressions in high resolution.
[1070] The device's microphone records the user's voice.
[1071] The captured data is sent to a server and analyzed by an emotion engine.
[1072] The analysis results, which represent the user's emotions (e.g., anxiety, fear, relief), are output to the server.
[1073] Step 5:
[1074] Calculation of overall score
[1075] The server calculates an overall score using the score calculation means based on the data obtained from the analysis means and emotion engine. The input is visitor attribute data and user emotion data, which are integrated to calculate and output an overall score.
[1076] Specific behavior:
[1077] The server combines visitor attribute data (e.g., age, gender, whether the face is familiar) with user emotion data (e.g., anxiety).
[1078] A score calculation means is used to quantify the overall score.
[1079] The calculated overall score is stored in a database on the server.
[1080] Step 6:
[1081] judgement
[1082] The server determines whether the visitor is suspicious or not based on the calculated total score. The input is the total score, and the output is the result of the determination whether the visitor is suspicious or not.
[1083] Specific behavior:
[1084] The server compares the overall score to a threshold.
[1085] If the score exceeds the threshold, the visitor is deemed suspicious.
[1086] The judgment results are stored in the server's database.
[1087] Step 7:
[1088] Notification and Response
[1089] If the visitor is determined to be suspicious, the device automatically responds to the visitor saying "We are unable to serve you," and the server sends a warning to the user's smartphone and even notifies the police. The input is the judgment result, and the output is the execution of the notification and report.
[1090] Specific behavior:
[1091] The device will send a voice notification to the visitor via the intercom saying "We are unavailable."
[1092] The server uses a service such as Twilio or Pushbullet to send an alert message to the user's smartphone.
[1093] The server notifies the police using a preset notification means.
[1094] (Application example 2)
[1095] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1096] In modern society, special frauds and other criminal activities are on the rise, posing a serious problem especially for the elderly and people living alone. Therefore, there is a need for advanced security systems that can efficiently analyze visitor identities and user emotions in real time to prevent criminal activities. However, conventional systems lack a comprehensive judgment function that takes into account not only visitor attributes but also user emotions, resulting in low fraud detection accuracy.
[1097] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1098] In this invention, the server includes: means for capturing video of visitors with a camera; means for capturing audio of visitors with a microphone; analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes; means for analyzing the user's video and audio to recognize the user's emotions; means for evaluating the visitor's attributes based on the extracted attributes and the recognized user emotions and calculating a score; means for determining whether the visitor is a suspicious person based on the calculated score; and means for notifying the user and sending a warning to the police if the visitor is determined to be a suspicious person. This realizes a comprehensive security system that integrates visitor identity analysis and user emotion recognition, making it possible to prevent specialized frauds and other criminal acts.
[1099] A "camera" is a device that captures video of visitors.
[1100] A "microphone" is a device that captures the visitor's voice.
[1101] A "generative artificial intelligence model" is a type of artificial intelligence that analyzes captured video and audio to extract visitor attributes.
[1102] "Analysis means" refers to means for analyzing captured video and audio, including generative artificial intelligence models.
[1103] A "user" is a person who uses this system.
[1104] The "means for recognizing emotions" is a means for identifying the emotions of a user by analyzing the user's video and audio.
[1105] The "means for evaluating attributes and calculating a score" is a means for evaluating attributes of a visitor based on the extracted attributes and the recognized user sentiment, and calculating a score.
[1106] The "means for determining whether a visitor is suspicious" is a means for determining whether a visitor is suspicious based on the calculated score.
[1107] The "means for notifying and sending a warning" is a means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[1108] This invention presents a form of advanced security system that analyzes visitor identity and user sentiment. The system consists of the following main components:
[1109] Hardware:
[1110] The system's terminal is a device equipped with a camera and microphone. This device functions as an intercom and captures real-time video and audio information of visitors, allowing the terminal to accurately record their actions and comments.
[1111] software:
[1112] The system applies a generative artificial intelligence model (AI model) and an emotion recognition engine that analyzes captured video and audio data to extract visitor attributes (age, gender, known or unknown) and user emotions (anxiety, relief, fear, etc.).
[1113] Analysis method:
[1114] Using AI models, the video and audio captured on the device are analyzed to extract visitor attributes.
[1115] Emotion recognition means:
[1116] The camera and microphone are used to analyze the user's facial expressions and tone of voice, and the emotion recognition engine identifies the user's emotions.
[1117] Data processing and calculation:
[1118] The server calculates a total score using the score calculation means based on the visitor's attributes extracted by the analysis means and the user's emotions identified by the emotion recognition means. This total score is used to determine whether the visitor is suspicious. Once a determination is made, the notification means is activated and a warning is sent to the user and the police.
[1119] Examples:
[1120] When a visitor standing at the front door operates the intercom, the camera automatically captures video and the microphone captures audio. This data is analyzed by an AI model to determine the visitor's age, gender, and whether they are a familiar face. Meanwhile, the user's camera video and audio are analyzed by an emotion recognition engine to identify the user's emotions. Based on this, the server calculates an overall score, integrating the visitor's attributes and the user's emotions to determine whether they are suspicious.
[1121] Example prompt sentence:
[1122] "Capture video of visitors and analyze their attributes."
[1123] "Estimate the visitor's age and gender and assess user sentiment."
[1124] "Use a comprehensive score to determine whether a visitor is suspicious."
[1125] In this way, the system provides comprehensive security that integrates visitor identity analysis and user emotion recognition, making it possible to effectively protect users from specialized fraud and other criminal activities.
[1126] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1127] Step 1:
[1128] The device uses a camera and microphone to capture the visitor's video and audio. The input is the visitor's real-time video and audio, and the output is the captured digital data. This digital data is temporarily stored for subsequent analysis processes.
[1129] Step 2:
[1130] The terminal sends the captured video and audio data to the analysis means. The input is the captured video and audio digital data, and the output is the visitor's attribute data (e.g., age, gender, whether the visitor is a known person) as the analysis result. The terminal analyzes the data using a generative artificial intelligence model and extracts the attributes.
[1131] Step 3:
[1132] The terminal captures the user's video and audio. This step is performed between the time the visitor operates the intercom and the time the user answers. The input is the user's real-time video and audio, and the output is the captured digital data.
[1133] Step 4:
[1134] The device sends the captured user video and audio data to the emotion recognition means. The input is the captured user video and audio digital data, and the output is the user's emotional data (e.g., anxiety, relief, fear, etc.) as an analysis result. The emotion recognition engine analyzes this data to identify the user's emotions.
[1135] Step 5:
[1136] The server integrates the output data of the analysis means and emotion recognition means and calculates an overall score using the score calculation means. The input is the visitor's attribute data and the user's emotion data, and the output is an overall score. The server performs data calculations based on this data and evaluates whether the visitor is suspicious or not.
[1137] Step 6:
[1138] The server determines whether the visitor is suspicious based on the total score. The input is the total score, and the output is the result of the judgment whether the visitor is suspicious. Based on this judgment result, a notification method is prepared.
[1139] Step 7:
[1140] If the server determines that a person is suspicious, it immediately notifies the user and sends a warning to the police. The input is the suspicious person determination result, and the output is a notification message to the user and a warning message to the police. The server automatically generates and sends the message.
[1141] In this way, the system analyzes the identity of visitors and the user's emotions in real time, enabling it to identify suspicious individuals and take prompt action.
[1142] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1143] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1144] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1145] [Fourth embodiment]
[1146] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1147] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1148] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1149] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1150] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1151] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1152] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1153] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1154] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1155] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1156] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1157] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1158] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1159] This invention is a system for preventing special fraud and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and aims to detect visitors who may be committing fraud.
[1160] System Configuration
[1161] The present invention comprises the following main components:
[1162] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[1163] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[1164] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[1165] 4. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates a score. The score is higher for elderly people, unfamiliar faces, and unnatural background sounds.
[1166] 5. The judgment means judges whether the visitor is suspicious or not based on the calculated score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[1167] 6. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[1168] System Operation
[1169] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. The server then evaluates the extracted attributes using a score calculation means to calculate a score. Based on this score, the server determines whether the visitor is suspicious.
[1170] As a specific example, the following scenario can be considered.
[1171] Example scenario:
[1172] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[1173] 2. The video shows a middle-aged man and the audio is mixed with unnatural noise. The video and audio data are automatically sent to an analysis device.
[1174] 3. The server analyzes the data, estimates the visitor's age to be 45, and verifies that the person has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background.
[1175] 4. Based on these attributes, the server calculates a score. For example, a high score can be set if the person's age is within a certain range, if they are not registered as an acquaintance, or if unnatural noise is detected.
[1176] 5. The visitor is determined to be suspicious based on the score calculated by the server.
[1177] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[1178] By using such a system, users can protect themselves from crime, and it can provide an effective crime prevention measure, especially for those who are more likely to become victims of crime, such as the elderly.
[1179] The processing flow will be explained below.
[1180] Step 1:
[1181] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[1182] Step 2:
[1183] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[1184] Step 3:
[1185] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[1186] Step 4:
[1187] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[1188] Step 5:
[1189] The server calculates a score for each visitor based on the extracted attributes. For example, an elderly person will receive a higher score, a familiar face will receive a lower score, and the presence of unnatural background noise will also result in a higher score.
[1190] Step 6:
[1191] The server uses the calculated score to determine whether the visitor is suspicious. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[1192] Step 7:
[1193] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[1194] Step 8:
[1195] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[1196] Step 9:
[1197] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[1198] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[1199] Example 1
[1200] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1201] In recent years, there has been an increase in special frauds and other crimes, with the elderly and other socially vulnerable people being particularly vulnerable. Conventional crime prevention measures monitor visitor information in real time, but it is often difficult to accurately identify suspicious individuals. Furthermore, ensuring the security of captured data and linking it to a prompt notification system for suspicious individuals have been issues.
[1202] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1203] In this invention, the server includes a means for capturing video of visitors using a camera, a means for capturing audio of visitors using a microphone, and a means for transmitting the captured and analyzed data using a secure protocol. This allows visitor information to be captured in real time, securely transmitted, and analyzed, making it possible to accurately evaluate visitor attributes and identify suspicious individuals. This also ensures data security and allows for prompt notification of suspicious individuals.
[1204] A "camera" is a device that captures video of visitors.
[1205] A "microphone" is a device that captures the visitor's voice.
[1206] A "generative artificial intelligence model" is an artificial intelligence model that analyzes captured video and audio data and extracts visitor attributes.
[1207] The "analysis means" is a means for analyzing the captured video and audio data to extract visitor attributes.
[1208] "Attributes" refers to characteristics such as the visitor's age, gender, whether they are a known person, and any abnormalities in background noise.
[1209] The "score calculation means" is a means for evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[1210] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[1211] The "notification means" is a means for notifying the user and the police if a person is determined to be a suspicious person.
[1212] A "secure protocol" is a communication method that ensures security in data transmission.
[1213] A "timestamp" is time information added to captured data.
[1214] A "data analysis module" is a piece of software functionality for extracting attributes from captured video and audio data.
[1215] A "score calculation module" is a part of the software function that calculates a score based on visitor attributes.
[1216] A "judgment module" is a part of the software function that judges whether a visitor is suspicious based on the calculated score.
[1217] This invention is a system for preventing special fraud and other crimes committed by visitors. The system is configured using the following hardware and software:
[1218] Hardware
[1219] 1. Camera: This is a device built into the intercom that captures video of visitors. When a visitor stands in front of the intercom, the camera automatically starts recording.
[1220] 2. Microphone: This is a device built into the intercom that captures the visitor's voice. When a visitor speaks, the microphone records the voice.
[1221] software
[1222] 1. Generative AI model: An AI model that analyzes captured video and audio data and extracts visitor attributes.
[1223] 2. Analysis method: A method for analyzing captured video and audio data to extract visitor attributes.
[1224] 3. Score calculation means: A means for evaluating the visitor's attributes based on the extracted attributes and calculating a score.
[1225] 4. Judgment method: A method for judging whether a visitor is suspicious or not based on the calculated score.
[1226] 5. Notification means: A means for notifying the user and the police based on the judgment results.
[1227] 6. Secure Protocol: A communication method that ensures security in data transmission.
[1228] 7. Data Analysis Module: This is the part of the software that functions to extract attributes from the captured video and audio data.
[1229] 8. Score calculation module: A part of the software function that calculates a score based on visitor attributes.
[1230] 9. Judgment module: This is the part of the software function that judges whether a visitor is suspicious or not based on the calculated score.
[1231] The system works as follows: First, the terminal captures the visitor's video and audio in real time using the intercom's built-in camera and microphone. Next, the captured data is sent to the server using a secure protocol (e.g., SSL / TLS). The server then analyzes the received video and audio data using a generative artificial intelligence model to extract attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. These attributes are evaluated using a score calculation means to calculate a score.
[1232] As a specific example, when a visitor presses the button on the intercom, the device activates the camera and microphone to capture the visitor's video and audio. The video shows a middle-aged man, and the audio is mixed with unnatural noise. The captured data is immediately sent to the server using a secure protocol, where it is analyzed. The analysis results indicate that the visitor is estimated to be 45 years old and that he has not previously been registered as an acquaintance. Audio analysis also detects unnatural noise in the background. Based on this information, the server calculates a high score and determines the visitor to be suspicious. Finally, a notification method is activated, sending an alert to the user's smartphone and also notifying the police.
[1233] Prompt Sentence Examples
[1234] "The video of a middle-aged man captured by an intercom camera contains unnatural noise. We will analyze this visitor to determine whether he is suspicious."
[1235] This system allows users to monitor visitor information in real time to ensure safety, and also provides effective crime prevention measures for vulnerable groups such as the elderly.
[1236] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1237] Step 1:
[1238] The device uses the intercom's camera and microphone to capture video and audio of visitors in real time. The camera automatically captures video when the visitor stands in front of the intercom, and the microphone records audio when the visitor speaks. The captured video and audio data are then time-stamped.
[1239] Input: Visitor video and audio
[1240] Output: Time-stamped video and audio data
[1241] How it works: When a visitor presses the button on the intercom, the camera and microphone automatically activate and record video and audio.
[1242] Step 2:
[1243] The device transmits the captured, time-stamped video and audio data to the server using a secure protocol (e.g., SSL / TLS), where the data is buffered and transmitted in real time.
[1244] Input: Time-stamped video and audio data
[1245] Output: Securely transmitted video and audio data
[1246] What happens: The device buffers the captured data and prepares to send it to the server over a secure channel. The transmission begins immediately.
[1247] Step 3:
[1248] The server uses a data analysis module to analyze the video and audio data it receives. First, video analysis involves facial recognition and age and gender estimation, while audio analysis involves the presence or absence of background noise and the characteristics of the audio.
[1249] Input: Securely transmitted video and audio data
[1250] Output: Visitor demographic data (e.g., age, gender, presence or absence of background noise)
[1251] How it works: The server analyzes the video data frame by frame to extract facial features, and simultaneously performs spectrogram analysis of the audio data to check for abnormal noise.
[1252] Step 4:
[1253] The server uses a score calculation module to calculate a score based on attributes extracted from the analysis results (e.g., the visitor's age, gender, whether they are a known person, whether there is background noise, etc.). The score is weighted for each attribute, and the overall likelihood of a suspicious person is evaluated.
[1254] Input: Visitor attribute data
[1255] Output: The calculated score
[1256] What happens: The server runs a program that estimates the visitor's age to be 45, determines their gender as male, checks for the presence of background noise, and calculates a score based on these attributes.
[1257] Step 5:
[1258] The server determines whether the visitor is suspicious based on the calculated score. If the visitor is determined to be suspicious, the server proceeds to the next notification process.
[1259] Input: Calculated score
[1260] Output: Result of suspicious person judgment (e.g., suspicious person / not suspicious person)
[1261] Specific operation: If the score exceeds a certain threshold, the server executes logic to determine the person as suspicious.
[1262] Step 6:
[1263] The notification mechanism will be activated, sending an alert to the user's smartphone and, if necessary, notifying the police, allowing the user to quickly understand the situation and take action.
[1264] Input: Suspicious person detection result
[1265] Output: Notification to user and police
[1266] Specific operation: The server sends a warning to the user's smartphone via SMS or app notification that there is a "suspicious person" and also notifies the police using an API.
[1267] (Application example 1)
[1268] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1269] Conventional visitor reception systems are good at collecting visitor information, but lack the ability to effectively analyze the collected information and quickly determine whether a visitor is suspicious. Furthermore, they lack the means to issue real-time warnings or reports, making it difficult to prevent crimes or take early action. The present invention aims to solve these problems by providing a system that prevents crimes committed by visitors and ensures the safety of users.
[1270] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1271] In this invention, the server includes means for capturing video of visitors with a camera, means for capturing audio of visitors with a microphone, analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes, means for evaluating the visitor attributes based on the extracted attributes and calculating a score, means for determining whether the visitor is a suspicious person based on the calculated score, means for notifying the user and sending a warning to the police if the visitor is determined to be suspicious, means for linking the intercom and smart device to transfer visitor information in real time, and means for analyzing visitor attributes in real time and calculating a score. This allows for quick and accurate analysis of visitor information, enabling the detection of suspicious people and immediate response.
[1272] A "camera" is a device that captures video of visitors.
[1273] A "microphone" is a device that captures the visitor's voice.
[1274] A "generative artificial intelligence model" is an analytical method that analyzes captured video and audio to extract visitor attributes.
[1275] The "score calculation means" is a means having a function of evaluating the attributes of a visitor based on the extracted attributes and calculating a score.
[1276] The "determination means" is a means for determining whether a visitor is suspicious or not based on the calculated score.
[1277] "Notification means" refers to the means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[1278] "Means for linking an intercom with a smart device" refers to a means for linking an intercom with a smart device such as a smartphone to transfer visitor information in real time.
[1279] The "real-time analysis means" is a means for analyzing visitor attributes in real time and calculating scores.
[1280] The system for implementing the present invention captures video and audio of visitors in real time and analyzes them using a generative artificial intelligence model to detect suspicious individuals and issue immediate warnings. This system uses the following hardware and software:
[1281] Hardware and Software Configuration
[1282] Hardware
[1283] Camera: A device that captures video of visitors, usually built into the intercom.
[1284] Microphone: A device that captures the visitor's voice and is built into the intercom, just like a camera.
[1285] Smart device: Usually a smartphone, which works in conjunction with the intercom to notify the user of visitor information.
[1286] software
[1287] Generative artificial intelligence model: An AI model that analyzes captured video and audio to extract visitor attributes (e.g., age, gender, background noise), for example using a deep learning framework such as Keras.
[1288] Score calculation algorithm: An algorithm for calculating the score based on the extracted attributes. It is implemented using a programming language such as Python.
[1289] Notification systems: Systems that send alerts to users when a visitor is deemed suspicious. This includes email sending services (such as smtplib) and mobile notification systems.
[1290] Process Overview
[1291] 1. The device uses a camera and microphone to capture video and audio of visitors in real time, and the captured data is immediately transmitted to an analytical tool that includes a generative artificial intelligence model.
[1292] 2. The server analyzes the video and audio data using a generative artificial intelligence model, which extracts attributes such as the visitor's age, gender, and the presence or absence of background noise.
[1293] 3. The server uses a scoring algorithm to calculate the visitor's score based on the extracted attributes, specifically, unknown visitors and those with unnatural noise are given a higher score.
[1294] 4. Based on the calculated score, the server determines whether the visitor is suspicious.
[1295] 5. If a score exceeds the threshold specified by the user, the system will identify the person as suspicious and immediately send a warning to the user's smart device. If requested, the system will also automatically notify the police.
[1296] Specific examples
[1297] For example, if a middle-aged man is seen on camera as a visitor and unnatural noise is detected in the background, the generative AI model will estimate his age to be 45 and detect background noise from the audio data. Based on this information, the scoring means will assign a high score and determine that the visitor is suspicious. The user will then receive a real-time notification saying, "A suspicious visitor has been detected!"
[1298] An example of an input prompt for a generative AI model is as follows:
[1299] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[1300] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1301] Step 1:
[1302] The device uses a camera and microphone to capture video and audio of visitors in real time.
[1303] Input: Real-time video data, real-time audio data
[1304] Data processing: Video and audio data is acquired and converted into digital format.
[1305] Output: Digital video data, digital audio data
[1306] Step 2:
[1307] The terminal transfers the captured video and audio data to the server.
[1308] Input: Digital video data, digital audio data
[1309] Data processing: Packetizing data and transferring it over the network
[1310] Output: Video and audio data transferred to the server
[1311] Step 3:
[1312] The server uses a generative artificial intelligence model to analyze the captured video and audio data.
[1313] Input: Video data, audio data
[1314] Data processing: Data analysis using generative AI models (extracting attributes such as age, gender, and background noise)
[1315] Output: Visitor demographic data (age, gender, presence of background noise, etc.)
[1316] Step 4:
[1317] The server uses a scoring algorithm to calculate a score for the visitor based on the extracted attributes.
[1318] Input: Visitor attribute data
[1319] Data processing: Score calculation algorithm (adds score according to attributes)
[1320] Output: Visitor score
[1321] Step 5:
[1322] The server determines whether the visitor is suspicious based on the calculated score.
[1323] Input: Visitor score
[1324] Data processing: Comparing scores with set thresholds
[1325] Output: Suspicious person determination result (suspicious person or non-suspicious person)
[1326] Step 6:
[1327] If a person is determined to be suspicious, the server will send a warning to the user's smart device and, if necessary, notify the police.
[1328] Input: Suspicious person judgment result
[1329] Data processing: generating and sending notification messages (contacting users and the police)
[1330] Output: Warning notice to user, report to police
[1331] Considering the following specific example, an example prompt would be:
[1332] "Please identify visitor attributes including: Age: 40-50, Gender: Male, Background Noise: Unnatural sounds."
[1333] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1334] This invention is a system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[1335] System Configuration
[1336] The present invention comprises the following main components:
[1337] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[1338] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[1339] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data to extract attributes such as the visitor's age, gender, whether the face is familiar, and any anomalies in the background sounds.
[1340] 4. The emotion engine is a means of recognizing the user's emotions, and uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions.
[1341] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[1342] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[1343] 7. The notification means notifies the user and also sends a warning to the police if a person is determined to be suspicious, allowing the user to take appropriate action promptly.
[1344] System Operation
[1345] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[1346] As a specific example, the following scenario can be considered.
[1347] Example scenario:
[1348] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[1349] 2. The video shows a middle-aged woman, and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice. In front of the intercom, the user's face takes on an anxious expression and her voice becomes trembling.
[1350] 3. The server analyzes the data, estimates the visitor's age to be 35, and confirms that the person has not previously been registered as an acquaintance. Furthermore, audio analysis detects no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[1351] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[1352] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[1353] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[1354] In this way, users can be more effectively protected from specialized fraud and other crimes by using a crime prevention system that combines emotion analysis, and the system can provide more advanced security measures to a wider range of people, including the elderly.
[1355] The processing flow will be explained below.
[1356] Step 1:
[1357] The device will activate the built-in camera and capture the visitor's video. When the visitor stands in front of the intercom, the video will be automatically captured.
[1358] Step 2:
[1359] The device uses a built-in microphone to capture the visitor's voice, and when the visitor speaks, the audio is recorded simultaneously.
[1360] Step 3:
[1361] The device transmits the captured video and audio data to an analysis means, which includes a generative artificial intelligence model that is used to analyze the data.
[1362] Step 4:
[1363] The server uses analytics to analyze the video and audio data to extract visitor attributes, including age, gender, whether the face is familiar, and any anomalies in the background sounds.
[1364] Step 5:
[1365] The device also captures the user's facial expressions and tone of voice to recognize their emotions. An emotion engine is used to analyze the user's emotions and identify emotions such as anxiety or fear.
[1366] Step 6:
[1367] The server calculates a score based on the extracted visitor attributes and the user's perceived emotions. For example, if an elderly person or unnatural noise is detected, the score will be higher. The score will also increase if the user shows signs of anxiety or fear.
[1368] Step 7:
[1369] The server determines whether the visitor is suspicious based on the calculated total score. If the score exceeds a certain threshold, the visitor is deemed suspicious.
[1370] Step 8:
[1371] If a visitor is deemed suspicious, the server will send an alert to the police, thus preventing crimes from occurring.
[1372] Step 9:
[1373] The device sends a notification to the user's smartphone, allowing the user to confirm that a suspicious person has visited the device and that the police have been notified.
[1374] Step 10:
[1375] Users are notified so they can be sure that they do not need to respond to the visitor and ensure their own safety.
[1376] These steps allow the system to determine in real time whether a visitor is suspicious and take appropriate action, thus protecting users from special fraud and other crimes.
[1377] Example 2
[1378] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1379] There is a need to quickly and accurately determine whether a visitor is suspicious and ensure the safety of users. Conventional systems only analyze visitor attributes and are unable to integrate information including user emotions. As a result, there is a problem that the accuracy of detecting suspicious individuals is low and users' anxiety cannot be completely alleviated.
[1380] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1381] In this invention, the server includes a means for capturing video of the visitor with a camera, a means for capturing the visitor's voice with a microphone, and a means for capturing the user's facial expression and tone of voice and analyzing the user's emotions, thereby making it possible to determine suspicious individuals by taking into account not only the visitor's attributes but also the user's emotions.
[1382] A "camera" is a device for capturing images.
[1383] A "microphone" is a device for capturing sound.
[1384] A "generative artificial intelligence model" is an artificial intelligence technology that analyzes specific features and patterns from input data and outputs the results.
[1385] "Analysis means" refers to a device or software for analyzing the captured video and audio data and extracting visitor attributes.
[1386] "Means for capturing a user's facial expression and tone of voice and analyzing the user's emotions" refers to a device or software that uses a camera and microphone to capture a user's facial expression and tone of voice and, based on that, identifies the user's emotions.
[1387] The "means for calculating a score" is a device or software for evaluating visitor attributes based on the extracted visitor attributes and the analyzed user sentiment.
[1388] The "means for determining" is a device or software for determining whether a visitor is suspicious or not based on the calculated score.
[1389] The "notification means" is a device or system that notifies the user when a person is determined to be suspicious and also sends a warning to the police.
[1390] This invention is a security system for preventing special frauds and other crimes committed by visitors. The system acquires visitor information using an intercom equipped with a camera and microphone, analyzes visitor attributes using generative artificial intelligence, and further recognizes user emotions using an emotion engine, with the aim of detecting visitors who may be committing fraud.
[1391] System Configuration
[1392] 1. The camera is a device that captures images of visitors and is built into the intercom. When a visitor stands in front of the intercom, the camera automatically takes an image.
[1393] 2. The microphone is a device that captures the visitor's voice and is built into the intercom just like the camera. When a visitor speaks, the microphone records the voice.
[1394] 3. The analysis means includes a generative artificial intelligence model for analyzing the captured video and audio data, which extracts attributes such as the visitor's age, gender, whether the face is familiar, and any abnormalities in the background sound. The generative artificial intelligence model used is a commonly used AI model (such as OpenAI's GPT-4).
[1395] 4. The emotion engine is a means of recognizing the user's emotions, and analyzes the user's facial expressions and tone of voice using a camera and microphone. Emotion engines that can be used include Amazon's AWS Rekognition and Google Cloud Vision API.
[1396] 5. The score calculation means evaluates the visitor's attributes based on the extracted attributes and calculates the overall score by correcting the score based on the recognized user's emotions. For example, if the user shows anxiety or fear, the score is set high.
[1397] 6. The judgment means judges whether the visitor is suspicious or not based on the calculated total score. If the visitor is judged to be suspicious, the process proceeds to the next step.
[1398] 7. If a person is determined to be suspicious, the notification method will notify the user and also send an alert to the police. This method allows the user to take appropriate action quickly. Services such as Twilio and Pushbullet can be used as notification methods.
[1399] System Operation
[1400] The device uses a camera and microphone to capture video and audio of the visitor in real time. The captured data is analyzed by an analysis means to extract the visitor's attributes. At the same time, the device captures the user's facial expressions and tone of voice, and the emotion engine analyzes the user's emotions. The server then uses a score calculation means to integrate the extracted visitor attributes and the user's emotions to calculate an overall score. Based on this score, the server determines whether the visitor is suspicious.
[1401] Specific examples
[1402] 1. The device captures the visitor's video using the intercom's camera and simultaneously captures the visitor's audio using the microphone.
[1403] 2. The video shows a middle-aged man and the audio includes clear speech. The device also captures the user's facial expressions and records the user's voice.
[1404] 3. The server analyzes the data, estimates the visitor's age to be 45, and confirms that this person has not previously been registered as an acquaintance. Furthermore, audio analysis detects that there is no abnormal background noise. Meanwhile, the emotion engine recognizes the user's anxious emotions.
[1405] 4. Based on these attributes and the user's emotions, the server calculates an overall score. For example, if the visitor is unknown and the user is anxious, the score is set high.
[1406] 5. The visitor is determined to be suspicious based on the overall score calculated by the server.
[1407] 6. The device will inform the visitor via the intercom that "we are unavailable," and the server will send an alert to the user's smartphone and then notify the police.
[1408] Examples of prompts:
[1409] Analyze the following data:
[1410] Visitor video data, age, gender, and whether they are known faces
[1411] Visitor voice data and whether there are any abnormal sounds in the background
[1412] Analysis results of user's facial expressions and tone of voice
[1413] The analysis methods are as follows:
[1414] 1. Identifying visitor attributes based on video data
[1415] 2. Analyze background sounds based on audio data
[1416] 3. Analyze user emotions
[1417] Calculate your overall score based on the analysis results.
[1418] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1419] Step 1:
[1420] Data Capture
[1421] The device captures the visitor's video and audio in real time. Specifically, the intercom captures video with a built-in camera and audio with a built-in microphone. The input is the visitor's real-time video and audio, and is used to output video and audio data.
[1422] Specific behavior:
[1423] The device's camera will then begin working, capturing high-resolution footage of the visitor.
[1424] The device's microphone will begin working and will record the visitor's speech and background sounds.
[1425] The captured video and audio data is temporarily stored in the device's memory.
[1426] Step 2:
[1427] Data transmission
[1428] The device sends the captured video and audio data to the server. The input is the captured video and audio data, which is then encrypted and sent to the server for output.
[1429] Specific behavior:
[1430] The terminal encrypts the video and audio data and transmits it to a server via a communications network.
[1431] Check whether the data was sent successfully and proceed to the next step.
[1432] Step 3:
[1433] Visitor attribute analysis
[1434] The server analyzes the received video data using a generative artificial intelligence model. The input is the video and audio data received from the device, and the output is the visitor's attribute data, such as age, gender, and whether or not the face is familiar.
[1435] Specific behavior:
[1436] The server inputs video data into the generative artificial intelligence model using prompt sentences and acquires attributes such as age and gender.
[1437] The server analyzes the audio data and determines whether there are any suspicious noises in the background sound.
[1438] The analysis results are stored in a database on the server.
[1439] Step 4:
[1440] User sentiment analysis
[1441] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and sends them to the server. The input is the user's facial expression data and tone of voice data, which is then used by the emotion engine to analyze the user's emotions and output them.
[1442] Specific behavior:
[1443] The device's camera captures the user's facial expressions in high resolution.
[1444] The device's microphone records the user's voice.
[1445] The captured data is sent to a server and analyzed by an emotion engine.
[1446] The analysis results, which represent the user's emotions (e.g., anxiety, fear, relief), are output to the server.
[1447] Step 5:
[1448] Calculation of overall score
[1449] The server calculates an overall score using the score calculation means based on the data obtained from the analysis means and emotion engine. The input is visitor attribute data and user emotion data, which are integrated to calculate and output an overall score.
[1450] Specific behavior:
[1451] The server combines visitor attribute data (e.g., age, gender, whether the face is familiar) with user emotion data (e.g., anxiety).
[1452] A score calculation means is used to quantify the overall score.
[1453] The calculated overall score is stored in a database on the server.
[1454] Step 6:
[1455] judgement
[1456] The server determines whether the visitor is suspicious or not based on the calculated total score. The input is the total score, and the output is the result of the determination whether the visitor is suspicious or not.
[1457] Specific behavior:
[1458] The server compares the overall score to a threshold.
[1459] If the score exceeds the threshold, the visitor is deemed suspicious.
[1460] The judgment results are stored in the server's database.
[1461] Step 7:
[1462] Notification and Response
[1463] If the visitor is determined to be suspicious, the device automatically responds to the visitor saying "We are unable to serve you," and the server sends a warning to the user's smartphone and even notifies the police. The input is the judgment result, and the output is the execution of the notification and report.
[1464] Specific behavior:
[1465] The device will send a voice notification to the visitor via the intercom saying "We are unavailable."
[1466] The server uses a service such as Twilio or Pushbullet to send an alert message to the user's smartphone.
[1467] The server notifies the police using a preset notification means.
[1468] (Application example 2)
[1469] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1470] In modern society, special frauds and other criminal activities are on the rise, posing a serious problem especially for the elderly and people living alone. Therefore, there is a need for advanced security systems that can efficiently analyze visitor identities and user emotions in real time to prevent criminal activities. However, conventional systems lack a comprehensive judgment function that takes into account not only visitor attributes but also user emotions, resulting in low fraud detection accuracy.
[1471] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1472] In this invention, the server includes: means for capturing video of visitors with a camera; means for capturing audio of visitors with a microphone; analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes; means for analyzing the user's video and audio to recognize the user's emotions; means for evaluating the visitor's attributes based on the extracted attributes and the recognized user emotions and calculating a score; means for determining whether the visitor is a suspicious person based on the calculated score; and means for notifying the user and sending a warning to the police if the visitor is determined to be a suspicious person. This realizes a comprehensive security system that integrates visitor identity analysis and user emotion recognition, making it possible to prevent specialized frauds and other criminal acts.
[1473] A "camera" is a device that captures video of visitors.
[1474] A "microphone" is a device that captures the visitor's voice.
[1475] A "generative artificial intelligence model" is a type of artificial intelligence that analyzes captured video and audio to extract visitor attributes.
[1476] "Analysis means" refers to means for analyzing captured video and audio, including generative artificial intelligence models.
[1477] A "user" is a person who uses this system.
[1478] The "means for recognizing emotions" is a means for identifying the emotions of a user by analyzing the user's video and audio.
[1479] The "means for evaluating attributes and calculating a score" is a means for evaluating attributes of a visitor based on the extracted attributes and the recognized user sentiment, and calculating a score.
[1480] The "means for determining whether a visitor is suspicious" is a means for determining whether a visitor is suspicious based on the calculated score.
[1481] The "means for notifying and sending a warning" is a means for notifying the user and sending a warning to the police if the user is determined to be a suspicious person.
[1482] This invention presents a form of advanced security system that analyzes visitor identity and user sentiment. The system consists of the following main components:
[1483] Hardware:
[1484] The system's terminal is a device equipped with a camera and microphone. This device functions as an intercom and captures real-time video and audio information of visitors, allowing the terminal to accurately record their actions and comments.
[1485] software:
[1486] The system applies a generative artificial intelligence model (AI model) and an emotion recognition engine that analyzes captured video and audio data to extract visitor attributes (age, gender, known or unknown) and user emotions (anxiety, relief, fear, etc.).
[1487] Analysis method:
[1488] Using AI models, the video and audio captured on the device are analyzed to extract visitor attributes.
[1489] Emotion recognition means:
[1490] The camera and microphone are used to analyze the user's facial expressions and tone of voice, and the emotion recognition engine identifies the user's emotions.
[1491] Data processing and calculation:
[1492] The server calculates a total score using the score calculation means based on the visitor's attributes extracted by the analysis means and the user's emotions identified by the emotion recognition means. This total score is used to determine whether the visitor is suspicious. Once a determination is made, the notification means is activated and a warning is sent to the user and the police.
[1493] Examples:
[1494] When a visitor standing at the front door operates the intercom, the camera automatically captures video and the microphone captures audio. This data is analyzed by an AI model to determine the visitor's age, gender, and whether they are a familiar face. Meanwhile, the user's camera video and audio are analyzed by an emotion recognition engine to identify the user's emotions. Based on this, the server calculates an overall score, integrating the visitor's attributes and the user's emotions to determine whether they are suspicious.
[1495] Example prompt sentence:
[1496] "Capture video of visitors and analyze their attributes."
[1497] "Estimate the visitor's age and gender and assess user sentiment."
[1498] "Use a comprehensive score to determine whether a visitor is suspicious."
[1499] In this way, the system provides comprehensive security that integrates visitor identity analysis and user emotion recognition, making it possible to effectively protect users from specialized fraud and other criminal activities.
[1500] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1501] Step 1:
[1502] The device uses a camera and microphone to capture the visitor's video and audio. The input is the visitor's real-time video and audio, and the output is the captured digital data. This digital data is temporarily stored for subsequent analysis processes.
[1503] Step 2:
[1504] The terminal sends the captured video and audio data to the analysis means. The input is the captured video and audio digital data, and the output is the visitor's attribute data (e.g., age, gender, whether the visitor is a known person) as the analysis result. The terminal analyzes the data using a generative artificial intelligence model and extracts the attributes.
[1505] Step 3:
[1506] The terminal captures the user's video and audio. This step is performed between the time the visitor operates the intercom and the time the user answers. The input is the user's real-time video and audio, and the output is the captured digital data.
[1507] Step 4:
[1508] The device sends the captured user video and audio data to the emotion recognition means. The input is the captured user video and audio digital data, and the output is the user's emotional data (e.g., anxiety, relief, fear, etc.) as an analysis result. The emotion recognition engine analyzes this data to identify the user's emotions.
[1509] Step 5:
[1510] The server integrates the output data of the analysis means and emotion recognition means and calculates an overall score using the score calculation means. The input is the visitor's attribute data and the user's emotion data, and the output is an overall score. The server performs data calculations based on this data and evaluates whether the visitor is suspicious or not.
[1511] Step 6:
[1512] The server determines whether the visitor is suspicious based on the total score. The input is the total score, and the output is the result of the judgment whether the visitor is suspicious. Based on this judgment result, a notification method is prepared.
[1513] Step 7:
[1514] If the server determines that a person is suspicious, it immediately notifies the user and sends a warning to the police. The input is the suspicious person determination result, and the output is a notification message to the user and a warning message to the police. The server automatically generates and sends the message.
[1515] In this way, the system analyzes the identity of visitors and the user's emotions in real time, enabling it to identify suspicious individuals and take prompt action.
[1516] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1517] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1518] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1519] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1520] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1521] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1522] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1523] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1524] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1525] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1526] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1527] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1528] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1529] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1530] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1531] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1532] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1533] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1534] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1535] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1536] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1537] The following is further disclosed regarding the above embodiment.
[1538] (Claim 1)
[1539] a means for capturing video of the visitor by a camera;
[1540] a means for capturing the visitor's voice by a microphone;
[1541] an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes;
[1542] means for evaluating visitor attributes based on the extracted attributes and calculating a score;
[1543] A means for determining whether a visitor is a suspicious person based on the calculated score;
[1544] means for notifying the user and sending a warning to the police when the user is determined to be a suspicious person;
[1545] A system including:
[1546] (Claim 2)
[1547] 10. The system of claim 1, further comprising means for determining the age and gender of the visitor based on the captured video.
[1548] (Claim 3)
[1549] 10. The system of claim 1, further comprising means for analyzing the captured audio to detect unnatural sound changes.
[1550] "Example 1"
[1551] (Claim 1)
[1552] a means for capturing video of the visitor by a camera;
[1553] a means for capturing the visitor's voice by a microphone;
[1554] an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes;
[1555] means for evaluating visitor attributes based on the extracted attributes and calculating a score;
[1556] A means for determining whether a visitor is a suspicious person based on the calculated score;
[1557] a means for notifying the user and the police according to the determination result;
[1558] means for transmitting the captured and analyzed data using a secure protocol;
[1559] A means of detecting and time-stamping anomalies in variables in real time;
[1560] A system including:
[1561] (Claim 2)
[1562] 10. The system of claim 1, further comprising means for estimating the visitor's age and gender based on the captured video.
[1563] (Claim 3)
[1564] 2. The system according to claim 1, further comprising means for analyzing the captured audio to detect unnatural changes in sound and reflecting the results in the attribute evaluation.
[1565] "Application Example 1"
[1566] (Claim 1)
[1567] a means for capturing video of the visitor by a camera;
[1568] a means for capturing the visitor's voice by a microphone;
[1569] an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes;
[1570] means for evaluating visitor attributes and calculating a score based on the extracted attributes;
[1571] A means for determining whether a visitor is a suspicious person based on the calculated score;
[1572] means for notifying the user and sending a warning to the police when the user is determined to be a suspicious person;
[1573] Linking the intercom with smart devices to transfer visitor information in real time,
[1574] A means to analyze visitor attributes in real time and calculate scores,
[1575] A system including:
[1576] (Claim 2)
[1577] 10. The system of claim 1, further comprising means for determining the age and gender of the visitor based on the captured video.
[1578] (Claim 3)
[1579] 10. The system of claim 1, further comprising means for analyzing the captured audio to detect unnatural sound changes.
[1580] "Example 2: Combining Emotion Engines"
[1581] (Claim 1)
[1582] a means for capturing video of the visitor by a camera;
[1583] a means for capturing the visitor's voice by a microphone;
[1584] an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes;
[1585] means for capturing a user's facial expression and tone of voice and analyzing the user's emotions;
[1586] means for evaluating visitor attributes based on the extracted attributes and the analyzed user sentiment and calculating a score;
[1587] A means for determining whether a visitor is a suspicious person based on the calculated score;
[1588] means for notifying the user and sending a warning to the police when the user is determined to be a suspicious person;
[1589] A system including:
[1590] (Claim 2)
[1591] 10. The system of claim 1, further comprising means for determining the age and gender of the visitor based on the captured video.
[1592] (Claim 3)
[1593] 10. The system of claim 1, further comprising means for analyzing the captured audio to detect unnatural sound changes.
[1594] "Application example 2 when combining emotion engines"
[1595] (Claim 1)
[1596] a means for capturing video of the visitor by a camera;
[1597] a means for capturing the visitor's voice by a microphone;
[1598] an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes;
[1599] means for analyzing the user's video and audio to recognize the user's emotions;
[1600] means for evaluating visitor attributes based on the extracted attributes and the recognized user sentiment and calculating a score;
[1601] A means for determining whether a visitor is a suspicious person based on the calculated score;
[1602] means for notifying the user and sending a warning to the police when the user is determined to be a suspicious person;
[1603] A system including:
[1604] (Claim 2)
[1605] 10. The system of claim 1, further comprising means for determining the age and gender of the visitor based on the captured video.
[1606] (Claim 3)
[1607] 10. The system of claim 1, further comprising means for analyzing the captured audio to detect unnatural sound changes. [Explanation of symbols]
[1608] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for capturing video of the visitor by a camera; a means for capturing the visitor's voice by a microphone; an analysis means including a generative artificial intelligence model that analyzes the captured video and audio to extract visitor attributes; means for evaluating visitor attributes based on the extracted attributes and calculating a score; A means for determining whether a visitor is a suspicious person based on the calculated score; means for notifying the user and sending a warning to the police when the user is determined to be a suspicious person; A system including:
2. 10. The system of claim 1, further comprising means for determining the age and gender of a visitor based on the captured video.
3. 10. The system of claim 1, further comprising means for analyzing the captured audio to detect unnatural sound changes.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A