System
The system addresses the challenge of managing suspicious visitors by automating intercom responses using a generative AI model for efficient and secure visitor management, reducing user burden and enhancing safety.
Patent Information
- Application Number
- JP2024116434
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
In modern society, there is an increasing risk from malicious door-to-door salesmen and suspicious individuals, particularly for elderly people and those living alone, who face challenges in managing visitor interactions, especially when recipients are not home, and determining whether a visitor is suspicious, with conventional intercom systems placing a heavy burden on users and lacking efficient automated responses.
A system that includes a signal receiving means for intercom signals, data analysis for identifying visitor information, a response generation using a generative AI model, response transmission, notification to a mobile device, user response acceptance, automatic response based on user absence mode, and learning and evaluation for suspicious person risk assessment, enabling safe and efficient visitor management.
The system automates visitor responses, reducing user burden and ensuring safe interactions by providing automated responses and risk assessments, allowing users to manage visitors efficiently and securely.
Smart Images

Figure 2026014960000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, risks from malicious door-to-door salesmen and suspicious individuals are increasing. Furthermore, the burden of dealing with visitors while working from home, and for elderly people and those living alone, is becoming a problem. It is particularly difficult to arrange redelivery when the recipient is not at home, and to determine whether a visitor is suspicious. To address these issues, there is a demand for safe and efficient visitor response services. [Means for solving the problem]
[0005] The present invention solves these problems with a system including a signal receiving means for receiving an intercom signal, a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information, a response generation means for generating a primary response using a generative AI model based on the analyzed data, a response transmission means for transmitting the generated response to the visitor via the intercom, a notification means for notifying a mobile device of the visitor information and the content of the generated response, a user response acceptance means for accepting a response decision from the mobile device, an automatic response means for generating an automatic response based on an away mode set by the user and transmitting it to the visitor, and a learning and evaluation means for learning visitor characteristics and comparing them with past data to evaluate the suspicious person risk. Furthermore, by further including a means for notifying a user of a high suspicious person risk by issuing a warning and issuing a shutout instruction as necessary when the learning and evaluation means evaluates the suspicious person risk as high, and a means for sending a shutout message to the visitor via the intercom, it is possible to reduce the burden of dealing with visitors and provide a safe living environment.
[0006] The "signal receiving means" is a device or component that has the function of receiving a signal emitted by the intercom and converting it into a format that can be processed within the system.
[0007] "Data analysis means" refers to an algorithm or system that analyzes the video and audio data obtained from the intercom and identifies visitor information.
[0008] A "response generation means" is a device or software that has the function of generating an appropriate primary response using a generative AI model based on analyzed data.
[0009] The "response transmitting means" is a device or system that has the function of transmitting the generated response to the visitor through the intercom.
[0010] The "notification means" is a device or system that has the function of transmitting visitor information and the generated response content to a mobile terminal.
[0011] The "user response receiving means" is a device or software that has the function of receiving a response decision from a mobile terminal and returning that information to the system.
[0012] An "automatic response means" is a device or software that has the functionality to automatically generate and send an appropriate response to a visitor based on the absence mode set by the user.
[0013] "Learning and evaluation tools" are algorithms and systems that learn visitor characteristics and compare them with past data to assess the risk of suspicious behavior.
[0014] The "shutout means" is a device or software that has the function of issuing a warning to the user when the learning and evaluation means evaluates that the visitor is at high risk of being a suspicious person, and if necessary, issuing an instruction to shut out the visitor. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention is a system for automating safe and efficient visitor attendance related to an intercom system. The system is implemented using the following components:
[0037] System configuration
[0038] 1. Signal receiving means:
[0039] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[0040] 2. Data analysis methods:
[0041] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[0042] 3. Response Generation Method:
[0043] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[0044] 4. Response sending method:
[0045] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[0046] 5. Means of notification:
[0047] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[0048] 6. User response acceptance means:
[0049] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether to respond to it themselves.
[0050] 7. Automated Response Methods:
[0051] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[0052] 8. Learning and assessment tools:
[0053] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[0054] 9. Shut-out measures:
[0055] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[0056] Explanation of program processing
[0057] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[0058] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[0059] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[0060] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[0061] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[0062] This system reduces the burden on users in dealing with visitors and allows them to communicate with them safely and efficiently.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] A visitor presses the intercom
[0066] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[0067] Step 2:
[0068] Sending data to the server
[0069] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[0070] Step 3:
[0071] First-order response by generative AI
[0072] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[0073] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[0074] The server uses the generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[0075] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[0076] Step 4:
[0077] Notification of visitor information to mobile app
[0078] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[0079] Step 5:
[0080] User response decision
[0081] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[0082] The user can then decide whether to respond based on that information.
[0083] Step 6:
[0084] User response
[0085] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[0086] The user can connect to the intercom using the app and begin talking directly with the visitor.
[0087] Step 7:
[0088] Out of Office Replies
[0089] The server will automatically generate the appropriate response if the user has configured unattended mode.
[0090] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[0091] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[0092] Step 8:
[0093] Learn visitor characteristics and identify suspicious individuals
[0094] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[0095] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[0096] If the risk score is high, the server will notify the user with a warning.
[0097] Step 9:
[0098] Shutdown execution
[0099] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[0100] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[0101] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[0102] Example 1
[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0104] In today's society, where safety and convenience are essential, the automation of visitor responses is becoming increasingly important. However, with conventional intercom systems, responses are performed manually, which places a burden on busy users. Furthermore, when a suspicious person visits, an immediate and appropriate response is required, but manual responses can be risky. Therefore, there is an urgent need to develop a system that automates visitor responses over intercoms in a safe and efficient manner.
[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0106] In this invention, the server includes receiving means for receiving intercom signals, analyzing means for analyzing video and audio data acquired from the intercom and identifying visitor information, generating means for generating a primary response using a generative AI model based on the analyzed data, transmitting means for sending the generated response to the visitor via the intercom, notifying means for notifying a mobile device of the visitor information and the generated response content, receiving means for accepting a response decision from the mobile device, automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, evaluation means for learning visitor characteristics and comparing them with past data to evaluate suspicious person risk, and means for notifying the mobile device of the evaluated visitor information. This enables automated visitor response, safe and efficient visitor management, and rapid response to suspicious person risks.
[0107] The "receiving means" is a device or program for receiving an intercom signal and acquiring video and audio data.
[0108] The "analysis means" is a device or program that has the function of analyzing the video and audio data acquired by the receiving means and identifying information about the visitor.
[0109] The "generation means" is a device or program that generates a primary response using a generative AI model based on the data analyzed by the analysis means.
[0110] The "transmitting means" is a device or program that transmits the response generated by the generating means to the visitor via the intercom.
[0111] The "notification means" is a device or program for notifying the mobile terminal of the visitor information and the generated response content.
[0112] The "accepting means" is a device or program that has the function of accepting a response decision from a mobile terminal.
[0113] An "automatic response means" is a device or program that has the function of generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[0114] The "assessment means" is a device or program that has the function of learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[0115] A "generative AI model" is an artificial intelligence algorithm or platform that generates appropriate responses based on analyzed data.
[0116] A "prompt" is an instruction given to a generative AI model that serves as a basis for generating a specific response.
[0117] The present invention provides an improved intercom system for automating visitor responses. The system includes a receiving unit, an analyzing unit, a generating unit, a transmitting unit, a notifying unit, a reception unit, an automatic response unit, and an evaluation unit.
[0118] Hardware and Software
[0119] 1. Receiving means
[0120] Hardware: Intercom, Server
[0121] Software: The receiving program receives signals from the intercom via the Internet and acquires video and audio data.
[0122] 2. Analysis method
[0123] Hardware: Server
[0124] Software: Uses voice recognition systems (e.g., Google Speech-to-Text) to convert voice data into text, and runs image processing algorithms (e.g., facial recognition technology) to analyze video data.
[0125] 3. Generation means
[0126] Hardware: Server
[0127] Software: Generative AI models (e.g., OpenAI GPT-3) are used to generate a first-order response based on the analyzed text and video data.
[0128] 4. Transmission Method
[0129] Hardware: Servers, intercoms
[0130] Software: Software that converts the response text into speech (e.g., a text-to-speech engine) and a transmitting program that sends it to the intercom.
[0131] 5. Means of notification
[0132] Hardware: Servers, mobile devices
[0133] Software: A program that uses the push notification function of the mobile app to notify the mobile device of visitor information and the generated response.
[0134] 6. Method of reception
[0135] Hardware: Mobile devices
[0136] Software: A mobile application that provides a user interface where users can view the response and choose whether to respond manually or via an automated response.
[0137] 7. Automated Response Methods
[0138] Hardware: Server
[0139] Software: A program that generates an automatic response message based on the absence mode set by the user, converts it into voice data, and sends it to the intercom.
[0140] 8. Evaluation Methods
[0141] Hardware: Server
[0142] Software: Machine learning algorithms that learn visitor characteristics and programs that assess risk of suspicious behavior based on historical data.
[0143] Specific processing flow
[0144] First, when a visitor presses the intercom button, the signal is sent to the server. The receiving means receives this and captures the video and audio data. Next, the server's analysis means converts the audio data into text using a voice recognition system, and analyzes the video data using facial recognition technology.
[0145] The generating means uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate primary response based on the text data and video data. This generated response is output as text data, converted into audio data by the transmitting means, and transmitted to the visitor via the intercom. Through this process, the visitor can receive the response generated by the AI.
[0146] The notification means also transmits the analysis results to the mobile device so that the user can check the response. The user can choose to respond manually using the reception means or leave it to the automatic response means. If the absence mode is set, the automatic response means automatically generates and transmits an appropriate response.
[0147] Furthermore, the evaluation means can learn visitor data, evaluate the risk of suspicious individuals, and issue a warning to the user as necessary.
[0148] Examples of specific examples and prompts
[0149] For example, if the visitor is a delivery person, the voice data of the delivery person saying "Delivery service" is sent to the server. The server uses a voice recognition system to generate the text "Delivery service," and the generative AI model then generates a response such as "Hello, you are a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom.
[0150] Examples of prompts include:
[0151] "Please tell me how you would respond if the visitor was a delivery person."
[0152] Please explain how to respond if a suspicious person visits.
[0153] This reduces the burden on users in dealing with visitors and allows them to communicate with visitors safely and efficiently.
[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0155] Step 1:
[0156] The visitor presses the intercom button.
[0157] Input: Visitor Action
[0158] Output: Signal from intercom to server
[0159] Specific operation: When the intercom button is pressed, video and audio data is generated and sent to the server along with a signal.
[0160] Step 2:
[0161] The server receives the signal transmitted from the interphone via the Internet.
[0162] Input: Intercom signal, video and audio data
[0163] Output: Video and audio data stored in a database
[0164] Specific operation: The server's receiving means captures the signal and stores the video and audio data in a database.
[0165] Step 3:
[0166] The server analyzes the received voice data and converts it into text using a voice recognition system (e.g., Google Speech-to-Text).
[0167] Input: Received audio data
[0168] Output: Text data
[0169] Specific operations: The server's analysis means analyzes the voice data using a voice recognition system and generates corresponding text data.
[0170] Step 4:
[0171] The server analyzes the received video data and attempts to identify the visitor using facial recognition technology.
[0172] Input: Received video data
[0173] Output: Visitor's specific information (face image, identification information)
[0174] What it does: The server's analytics runs the video data through image processing algorithms to recognize and identify the visitor's face and stores that information.
[0175] Step 5:
[0176] The server generates a first-order response based on the analyzed text and video data using a generative AI model (e.g., OpenAI GPT-3).
[0177] Input: Text data, video data
[0178] Output: Text data of the primary responses
[0179] Specific operation: The server's generation means inputs text data and video data into the generative AI model and generates a first response such as "Hello, how can I help you?"
[0180] Step 6:
[0181] The server converts the generated primary response into voice data and transmits it to the intercom.
[0182] Input: Text data of the primary response
[0183] Output: Audio data, sent to intercom
[0184] Specific operation: The server's transmission means converts the primary response into voice data using a text-to-speech engine, and transmits the voice data to the intercom.
[0185] Step 7:
[0186] The server notifies the mobile terminal of the analysis results and the generated response.
[0187] Input: Visitor information, content of initial response
[0188] Output: Notification message to mobile device
[0189] Specific operation: The server's notification means uses a push notification service to send the visitor's video, part of the audio text, and the response content to the mobile device.
[0190] Step 8:
[0191] The mobile terminal provides an interface for the user to check the notification content and select a manual or automatic response.
[0192] Input: Notification message from the server
[0193] Output: User response selection (manual or automatic)
[0194] What it does: Displays an interface on the mobile device that allows the user to view the notification and then provide options for how to respond.
[0195] Step 9:
[0196] If the user has enabled away mode, the server automatically generates an appropriate response and sends it to the visitor.
[0197] Input: User's away mode setting
[0198] Output: Auto-answer message, sent to intercom
[0199] Specific operation: The server's automatic response means uses the generative AI model to generate a message such as "I'm not here right now," converts it into voice data, and sends it to the intercom.
[0200] Step 10:
[0201] The server learns the visitor's characteristic data and compares it with past data to assess the risk of suspicious behavior.
[0202] Input: Visitor video and audio data, past visitor data
[0203] Output: Risk assessment results, warning notifications if necessary
[0204] Specific operation: The server's evaluation means uses a machine learning algorithm to analyze visitor data, assess the risk of suspicious activity, and if the risk is high, issue a warning to the user.
[0205] This detailed processing step ensures visitor interaction is automated, secure, and efficient.
[0206] (Application example 1)
[0207] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0208] With conventional intercom systems, communication with visitors is done manually, which places a heavy burden on the person in charge and makes it difficult to respond, especially when the person in charge is not present. Furthermore, in busy environments such as logistics centers, efficient visitor response is required. Furthermore, it is difficult to assess the risk of suspicious individuals, and measures to ensure safety are insufficient. To solve these issues, a system that automates visitor response safely and efficiently is needed.
[0209] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0210] In this invention, the server includes a signal receiving means for receiving an intercom signal, a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information, a response generation means for generating a primary response using a generative AI model based on the analyzed data, a response transmission means for sending the generated response to the visitor via the intercom, a notification means for notifying a mobile device of the visitor information and the generated response, a user response receiving means for receiving a response decision from the mobile device, an automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, a learning and evaluation means for learning visitor characteristics and comparing them with past data to evaluate the risk of suspicious activity, a means for converting visitor information into text using a voice recognition system, a means for generating a primary response based on the text data using a generative AI model, a means for converting the generated text response into voice data, a communication means for notifying a smartphone or tablet of the analyzed visitor information, and a means for providing a user interface for determining a response based on the notified information. This reduces the burden on visitors and enables safe and efficient visitor response.
[0211] The "signal receiving means" is a means for receiving a signal from the interphone and processing the signal.
[0212] The "data analysis means" is a means for analyzing the video and audio data acquired from the intercom and identifying the visitor's information.
[0213] A "response generation means" is a means for generating a primary response using a generative AI model based on the analyzed data.
[0214] The "response transmitting means" is a means for transmitting the generated response to the visitor via the intercom.
[0215] The "notification means" is a means for notifying the mobile terminal of the visitor information and the generated response content.
[0216] The "user response receiving means" is a means for receiving a response decision from a mobile terminal.
[0217] The "automatic response means" is a means for generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[0218] The "learning and evaluation means" is a means for learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[0219] A "voice recognition system" is a system for converting visitor information from voice data into text.
[0220] A "generative AI model" is an artificial intelligence model for generating first-order responses based on text data.
[0221] "Speech conversion means" refers to means for converting the generated text response into voice data.
[0222] "Communication means" refers to the means for notifying analyzed visitor information to a smartphone or tablet.
[0223] A "user interface" is a means for providing an interface for determining a response based on notified information.
[0224] To put the present invention into practice, the following system and its operation will be described. This system automates visitor reception at a logistics center in a safe and efficient manner.
[0225] First, a signal receiving means is provided to receive an intercom signal. This signal receiving means serves to acquire video and audio data generated by pressing the intercom button. Next, a data analysis means analyzes the video and audio data acquired from the intercom and identifies visitor information. The identified information is processed by a response generation means that generates a primary response using a generative AI model based on the analyzed data.
[0226] The generated primary response is sent to the visitor via the intercom via the response sending means. At the same time, the visitor information and the generated response are sent to the administrator's mobile device via the notification means. The mobile device is assumed to be a smartphone or tablet.
[0227] The mobile terminal is provided with a user response reception means, and the administrator can decide whether to respond manually or continue with an automatic response based on the received notification. The response decision may also be made automatically based on the absence mode set by the user. In this case, the automatic response means generates an appropriate response and sends it to the visitor.
[0228] In addition, the server has a learning and evaluation means that learns the characteristics of visitors and compares them with past data to evaluate the risk of suspicious persons. If the risk of suspicious persons is evaluated as high, a warning is sent to the user and, if necessary, a shut-out instruction is executed. A shut-out message is sent to the visitor via the intercom.
[0229] The system also includes a means for converting visitor information into text using a speech recognition system. For example, the Google Cloud Speech-to-Text API is used to convert the visitor's speech into text. A primary response based on the text data is generated using a generative AI model (e.g., OpenAI GPT-4). The generated text response is converted into voice data using a speech conversion means (e.g., Amazon Polly).
[0230] The server also has a communication method for notifying the smartphone or tablet of the analyzed visitor information. This communication is performed using a service such as Firebase Cloud Messaging (FCM). A UI framework such as React Native is used to provide a user interface for determining a response based on the notified information.
[0231] For example, if a visitor says, "Today's delivery," this voice data is converted into text data, "Today's delivery," via the Google Cloud Speech-to-Text API. A generative AI model receives this text data and generates a response, for example, "Hello, you're the delivery person. Please wait a moment." This response is converted into voice data by Amazon Polly and conveyed to the visitor via the intercom. At the same time, this information is also notified to the administrator's smartphone.
[0232] Example prompt: "Generate an appropriate response if the visitor says, 'Delivery today.'"
[0233] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0234] Step 1:
[0235] When the interphone button is pressed, the signal receiving means receives a signal from the interphone. This signal includes video data and audio data. After receiving the signal, these data are transmitted to the server.
[0236] Input: Intercom signal
[0237] Output: Video data, audio data
[0238] Step 2:
[0239] The server analyzes the received video and audio data using a data analysis tool. The audio data is converted to text using the Google Cloud Speech-to-Text API. The video data is analyzed using a facial recognition algorithm to obtain visitor information.
[0240] Input: Video data, audio data
[0241] Output: Text data, visitor information (face recognition results)
[0242] Step 3:
[0243] The server generates a first response using a generative AI model based on the analyzed text data and visitor information. OpenAI GPT-4 is used for this. The generated first response is output as text data.
[0244] Input: Text data, visitor information
[0245] Output: Text data of the primary response
[0246] Step 4:
[0247] The server converts the generated text data of the primary response into voice data using Amazon Polly as a voice conversion means.
[0248] Input: Text data of the primary response
[0249] Output: Audio data of the first response
[0250] Step 5:
[0251] The server transmits the generated voice data of the primary response to the intercom through the response transmitting means, thereby notifying the visitor of the response.
[0252] Input: Primary response audio data
[0253] Output: Visitor receives a voice response
[0254] Step 6:
[0255] The server notifies the administrator of the analyzed visitor information and the generated primary response content via a notification means to the administrator's smartphone or tablet.
[0256] Input: Visitor information, temporary response text data
[0257] Output: Notification to the administrator's smartphone
[0258] Step 7:
[0259] When the administrator receives a notification on their smartphone or tablet, the visitor's video, audio, and initial response are displayed, allowing the user to decide whether to respond manually or select an automatic response.
[0260] Input: Notified information (visitor's video, audio text, initial response)
[0261] Output: User response prompt
[0262] Step 8:
[0263] If the user judges that the visitor is a high risk of being a suspicious person, the server uses learning and evaluation means to evaluate the risk, and if necessary, issues a warning to the administrator and executes a shut-out instruction. A shut-out message is sent to the visitor via the intercom.
[0264] Input: Information about suspicious person risk, user shut-out instructions
[0265] Output: Shutout message
[0266] Step 9:
[0267] The server analyzes and learns from visitor characteristics and past data via learning and evaluation means, and reflects this in future responses, thereby improving the accuracy of future visitor risk assessments.
[0268] Input: Visitor characteristics, historical data
[0269] Output: Learning results, updated evaluation model
[0270] In this way, visitor reception can be automated, safely, and efficiently achieved.
[0271] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0272] The present invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components.
[0273] System configuration
[0274] 1. Signal receiving means:
[0275] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[0276] 2. Data analysis methods:
[0277] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[0278] 3. Response Generation Method:
[0279] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[0280] 4. Response sending method:
[0281] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[0282] 5. Means of notification:
[0283] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[0284] 6. User response acceptance means:
[0285] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether or not to respond to it.
[0286] 7. Automated Response Methods:
[0287] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[0288] 8. Learning and assessment tools:
[0289] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[0290] 9. Shut-out measures:
[0291] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[0292] 10. Emotion Engine:
[0293] The server is equipped with an emotion engine that analyzes the user's video and audio data to recognize the user's emotions. The emotion engine analyzes the emotions (e.g., stress, anxiety, joy, surprise, etc.) that the user shows while interacting with visitors and adjusts the system's response based on that information.
[0294] 11. Emotion-based response adjustment measures:
[0295] The server adjusts the operation of the notification means and response generation means based on the user's emotions recognized by the emotion engine. For example, if the user is feeling stressed or anxious, the response generation means generates a response with a gentler tone or switches to an automatic response.
[0296] Explanation of program processing
[0297] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[0298] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[0299] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[0300] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[0301] Furthermore, the server analyzes the user's video and audio using an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed, the emotion engine detects this and the server adjusts the response generation means based on that information. As a result, the generated response is delivered in an appropriate tone according to the user's emotional state.
[0302] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[0303] This system reduces the burden on users when dealing with visitors and allows them to communicate with them safely and efficiently. In addition, the introduction of an emotion engine enables responses that take into account the user's emotional state, providing even greater peace of mind and convenience.
[0304] The processing flow will be explained below.
[0305] Step 1:
[0306] A visitor presses the intercom
[0307] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[0308] Step 2:
[0309] Sending data to the server
[0310] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[0311] Step 3:
[0312] First-order response by generative AI
[0313] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[0314] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[0315] The server uses a generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[0316] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[0317] Step 4:
[0318] Notification of visitor information to mobile app
[0319] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[0320] Step 5:
[0321] User response decision
[0322] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[0323] The user can then decide whether to respond based on that information.
[0324] Step 6:
[0325] User response
[0326] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[0327] The user can connect to the intercom using the app and begin talking directly with the visitor.
[0328] Step 7:
[0329] Out of Office Replies
[0330] The server will automatically generate the appropriate response if the user has configured unattended mode.
[0331] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[0332] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[0333] Step 8:
[0334] Learn visitor characteristics and identify suspicious individuals
[0335] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[0336] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[0337] If the risk score is high, the server will notify the user with a warning.
[0338] Step 9:
[0339] Shutdown execution
[0340] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[0341] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[0342] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[0343] Step 10:
[0344] Recognizing user emotions with an emotion engine
[0345] The server acquires the user's video and audio data from the user's mobile terminal.
[0346] The server uses an emotion engine to analyze the user's video and audio data and recognize the user's emotions.
[0347] Step 11:
[0348] Emotion-Based Response Modulation
[0349] The server adjusts the response generation means based on the user's emotions recognized by the emotion engine.
[0350] If the user shows signs of stress or anxiety, the server generates a response in a gentle tone using the response generating means, or switches to an automatic response.
[0351] Example 2
[0352] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0353] Conventional intercom systems require users to respond to visitors manually, placing a heavy burden on users. It is also difficult to respond appropriately when the resident is away or to dangerous visitors, resulting in safety and convenience issues. Furthermore, the system is unable to respond flexibly to the user's emotional state, which can be stressful.
[0354] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a signal receiving means for receiving an intercom signal; a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information; a response generation means for generating a primary response using a generative model based on the analyzed data; a response transmission means for transmitting the generated response to the visitor via the intercom; a notification means for notifying the mobile communication device of the visitor information and the generated response content; a user response receiving means for receiving a response decision from the mobile communication device; an automatic response means for generating an automatic response based on an absence mode set by the user and transmitting it to the visitor; a learning and evaluation means for learning visitor characteristics and evaluating the risk of a suspicious person by comparing them with past data; a means for analyzing the user's video and audio data and having an emotion engine that recognizes the user's emotions; and an emotion-based response adjustment means for adjusting the operation of the notification means and the response generation means based on the recognized user emotions. This automates visitor response, reducing the burden on the user and enabling appropriate responses to be taken when the user is absent or when a dangerous visitor is present. Furthermore, a system that can respond flexibly according to the user's emotional state is provided, allowing the system to be used with peace of mind.
[0355] The "signal receiving means" is a device or function for receiving a signal transmitted from the intercom.
[0356] "Data analysis means" refers to a device or function that analyzes the video and audio data acquired from the intercom and identifies visitor information.
[0357] A "response generator" is a device or function that uses a generative model based on analyzed data to generate a primary response.
[0358] The "response transmitting means" is a device or function for transmitting the generated response to the visitor via the intercom.
[0359] The "notification means" refers to a device or function for notifying the portable communication device of the visitor's information and the generated response content.
[0360] The "user response receiving means" is a device or function for receiving a response decision from the portable communication device.
[0361] The "automatic response means" is a device or function that generates an automatic response based on the absence mode set by the user and sends it to the visitor.
[0362] "Learning and evaluation means" refers to devices and functions that learn the characteristics of visitors and compare them with past data to evaluate the risk of suspicious persons.
[0363] An "emotion engine" is a device or function that analyzes a user's video and audio data and recognizes the user's emotions.
[0364] The "emotion-based response adjustment means" is a device or function for adjusting the operation of the notification means or response generation means based on the recognized user emotion.
[0365] This invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining it with an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components:
[0366] System configuration
[0367] 1. Signal receiving means:
[0368] The server receives the intercom signal and captures the video and audio data. When a visitor presses the intercom, the signal is sent to the server via the Internet. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server.
[0369] 2. Data analysis methods:
[0370] The server receives the video and audio data sent from the intercom. It converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text, and then compares it with the video data. For example, a voice saying "This is a delivery" can be converted into text like "Takhaibindesu."
[0371] 3. Response Generation Method:
[0372] The server generates an appropriate response using a generative AI model (e.g., OpenAI's GPT-4) based on the text and video data from the voice. The generated response is first output in text format and then reconverted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[0373] 4. Response sending method:
[0374] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[0375] 5. Means of notification:
[0376] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The notification includes an image of the visitor's face, a portion of the voice text, and the response. For example, the user's smartphone may receive a notification saying, "A delivery person has visited you. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[0377] 6. User response acceptance means:
[0378] The mobile device displays the received notification for the user to review. The user can then review the notification and choose to respond to it themselves, have it automatically respond, or shut it out. For example, if the user is busy, they can choose to automatically respond.
[0379] 7. Automated Response Methods:
[0380] If the user is in away mode, the server automatically generates an appropriate response to send to the intercom, such as "I'm not here right now. Please come back later."
[0381] 8. Learning and assessment tools:
[0382] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, it sends a warning to the user.
[0383] 9. Shut-out measures:
[0384] Users can select the shut-out option on their mobile app, in which case the server will generate a shut-out message and send it to the visitor over the intercom, such as "Please refrain from visiting."
[0385] 10. Emotion Engine:
[0386] The server is equipped with an emotion engine that analyzes the user's video and audio data and recognizes the user's emotions, such as stress, anxiety, joy, surprise, etc.
[0387] 11. Emotion-based response adjustment measures:
[0388] The server adjusts the content and tone of the generative AI model's responses based on the user's perceived emotions. For example, if the user is feeling stressed, the server can generate responses in a gentler tone or switch to an automatic response mode.
[0389] Examples of specific examples and prompts
[0390] For example, consider a scenario where a delivery person arrives.
[0391] 1. Signal reception: The delivery person presses the intercom, and a "ping pong" signal is sent to the server.
[0392] 2. Analysis of audio and video data: The server converts the audio "This is a delivery" into the text "Takhaibindesu" and recognizes the visitor's face.
[0393] 3. Generate a response: The generative AI model (e.g., GPT-4) generates a response such as, "Hello, you're a delivery person. Please wait a moment."
[0394] 4. Sending response: The server converts the response into voice data and sends it back to the intercom to be conveyed to the delivery person.
[0395] 5. User notification: The server notifies the user of the visitor information and the response to their mobile device. The user receives a notification that a delivery person has visited them. Response: "Hello, it's you, a delivery person. Please wait a moment."
[0396] 6. Accepting user response: The user opens the app and selects how to respond. Since they are busy, they select automatic response.
[0397] 7. Execute Auto-Reply (Away Mode): If the user has Away Mode enabled, a message will be generated saying "I'm out of the office right now. Please come back later."
[0398] 8. Learning and risk assessment: The server learns from the video and audio of the delivery person and performs risk assessment. It determines that the person is not suspicious.
[0399] 9. Execute Shutout: Send a shutout message depending on conditions such as absence or risk of suspicious person. Not necessary this time.
[0400] 10. Analysis by emotion engine: The emotions (e.g., anxiety) that users feel when interacting with visitors are analyzed.
[0401] 11. Adjusting responses based on emotion: If the server is stressed, it will generate a gentler response such as "Please wait a moment."
[0402] For example, the prompt to input to a generative AI model might look like this:
[0403] Caller's voice: "This is a courier."
[0404] Prompt to generative AI: "The visitor says 'This is a courier.' Please generate an appropriate first-level response."
[0405] This allows the system to provide highly efficient visitor support.
[0406] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0407] Step 1: Signal reception
[0408] When a visitor presses the intercom, the server receives the signal via the Internet. This signal contains the visitor's video and audio data. The input is the signal when the visitor presses the intercom, and the output is the video and audio data stored on the server. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server, and the server captures the video and audio at that moment.
[0409] Step 2: Data analysis
[0410] The server analyzes the received video and audio data. First, it converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text. Next, it analyzes the video data and recognizes the visitor's face. This converts the input audio data into text format and also generates facial recognition data. For example, the audio saying "This is a delivery person" is converted into the text "Takhaibindesu" and the visitor's face is identified.
[0411] Step 3: Response Generation
[0412] The server generates an appropriate first-order response using a generative AI model (e.g., OpenAI's GPT-4) based on the text data and facial recognition data. The input is the text-converted voice data and facial recognition data, and the output is the response text. This response text is then converted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[0413] Step 4: Send response
[0414] The server sends the generated voice data to the intercom, and the visitor can receive the generated response. The input is the voice data generated by the server, and the output is the voice response transmitted to the visitor through the intercom. For example, the voice transmitted to the delivery person might say, "Hello, you are a delivery person. Please wait a moment."
[0415] Step 5: Notification
[0416] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The input is the visitor information (face image, part of the voice text) and the generated response, and the output is a notification message. A notification containing the visitor's face image and text is displayed on the user's mobile device. For example, the notification may read, "A delivery person has visited. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[0417] Step 6: Accept user response
[0418] The device displays the received notification for the user to review. The user can then review it and choose to respond manually, have it automatically respond, or shut it out. The input is the notification message from the server, and the output is the user's response decision. For example, if the user is busy and chooses to have it automatically respond, they can select it through the app.
[0419] Step 7: Auto-responders
[0420] If the user is in away mode, the server automatically generates an appropriate response and sends it to the intercom. The input is the away mode setting, and the output is an automatic response message. For example, a message like "I'm not available right now. Please come back later" is generated.
[0421] Step 8: Learning and risk assessment
[0422] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. The input is past visitor data and current visitor data, and the output is the suspicious person risk assessment result. If the server determines there is a high risk, it sends a warning to the user.
[0423] Step 9: Shut Out
[0424] The user can select the shut-out option on the mobile app. In this case, the server generates a shut-out message and transmits it to the visitor through the intercom. The input is the user's shut-out instruction, and the output is the shut-out message. For example, a message such as "Please refrain from visiting" is generated.
[0425] Step 10: Emotion Engine Analysis
[0426] The server analyzes the user's video and audio data and recognizes the user's emotions. The input is the user's video and audio data, and the output is the recognized emotional information. For example, it may be recognized from the user's video that they are feeling stressed.
[0427] Step 11: Adjust your response based on your emotions
[0428] The server adjusts the behavior of the notification and response generation means based on the recognized user emotion. The input is the recognized emotion information, and the output is a response adjusted according to the emotion. For example, if the user is feeling stressed, the generative AI model can generate a gentler tone of response or switch to automatic response mode.
[0429] (Application example 2)
[0430] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0431] With conventional intercom systems, responding to visitors is done manually, which is not only time-consuming for users but also poses issues in terms of safety and efficiency. Even with electronic payment services, when a problem occurs during a transaction, the response is often delayed, causing frustration for users. Furthermore, insufficient evaluation of fraud risk leaves users with security concerns.
[0432] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means. This automates visitor response and transaction management, enabling safe and efficient responses while taking into consideration the user's emotional state.
[0433] The "signal receiving means" is a function that allows the server to receive signals from the interface system.
[0434] The "data analysis means" is a function for analyzing the video and audio data acquired from the interface system to identify user information.
[0435] The "response generation means" is a function for generating a primary response using a generative AI model based on the analyzed data.
[0436] The "response sending means" is a function for sending the generated response to the user through the interface system.
[0437] The "notification means" is a function for notifying the mobile terminal of the user information and the generated response content.
[0438] The "user response receiving means" is a function for receiving a response decision from a mobile terminal.
[0439] The "automatic response means" is a function for generating an automatic response based on the absence mode set by the user and sending it to the user.
[0440] The "learning and evaluation means" is a function for learning from past transaction data and evaluating fraud risk.
[0441] The "emotion analysis means" is a function for analyzing the emotional state of the user.
[0442] The "response adjustment means" is a function for adjusting the response based on the analysis results of the emotion analysis means.
[0443] The present invention provides a system for automating visitor response and transaction management, enabling safe and efficient response while taking into consideration the emotional state of the user. The system includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means.
[0444] System configuration and operation
[0445] 1. Signal receiving means:
[0446] The server receives a signal from an interface system (e.g., an electronic payment terminal) to obtain transaction data.
[0447] 2. Data analysis methods:
[0448] The server analyzes the video and audio data acquired from the interface system to identify the user's information. Specifically, it converts the audio data into text using a speech recognition system (e.g., Google Cloud Speech-to-Text) and analyzes it.
[0449] 3. Response Generation Method:
[0450] A generative AI model (e.g., GPT-4) is used to generate a first-order response based on the analyzed data. Example prompts for response generation are as follows:
[0451] "User is experiencing payment error. Stressed. Please generate an appropriate customer support response."
[0452] 4. Response sending method:
[0453] The generated response is converted into text or speech and sent to the user through the interface system, using text-to-speech (TTS) technology (e.g., Amazon Polly) to reconvert the speech.
[0454] 5. Means of notification:
[0455] The server then sends the user information and the generated response to the mobile device (smartphone) using a real-time notification system such as Firebase Cloud Messaging.
[0456] 6. User response acceptance means:
[0457] The mobile terminal provides an interface for receiving the notification and accepting the user's response. The user can respond manually or select an automatic response.
[0458] 7. Automated Response Methods:
[0459] Based on the out-of-office mode set by the user, an automated response is generated and sent to the user. The out-of-office response is constructed by the generative AI and includes a message informing the user that they are out of the office.
[0460] 8. Learning and assessment tools:
[0461] The server has an algorithm that learns from past transaction data and evaluates fraud risk. If the fraud risk is assessed as high, it will send a warning to the user.
[0462] 9. Emotion analysis means:
[0463] The server uses an emotion engine (e.g., IBM Watson or Microsoft Azure Emotion API) to analyze the user's emotional state, detecting emotions such as stress, anxiety, and joy.
[0464] 10. Response Adjustment Measures:
[0465] The server adjusts the response based on the analysis results of the emotion analysis means, and if the user is feeling stressed, it generates a response with a more reassuring tone.
[0466] Examples:
[0467] For example, consider a scenario in which a user encounters a payment error while using an electronic payment service and seeks support. A signal is sent from the interface system to the server, which converts the voice data into text. The server generates an appropriate response, converts it into voice, and sends it to the user. At the same time, the mobile device is notified of the details of the problem and the generated response. Furthermore, the server uses an emotion engine to analyze the user's emotional state and adjust the response accordingly. An example of a specific prompt would be, "The user is experiencing a payment error. He is in a stressful state. Please generate an appropriate customer support response."
[0468] This system reduces the burden on users in dealing with visitors and managing transactions, allowing them to receive safe and efficient services. In addition, the introduction of an emotion engine makes it possible to respond according to the user's emotional state, providing even greater peace of mind and convenience.
[0469] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0470] Step 1:
[0471] Signal Reception
[0472] The server receives signals from the interface system (e.g., electronic payment terminal). These signals include visitor information and transaction data. The input is the signal from the interface system, and the output is the received signal data. The signal sent from the interface system is received by the server's API.
[0473] Step 2:
[0474] Data analysis
[0475] The server analyzes the received video and audio data. The audio data is first converted into text using a speech recognition system (e.g., Google Cloud Speech-to-Text). Analysis is performed based on this text data and video data. The input is the received signal data, and the output is text data and analyzed video data. The audio data is converted into text format and then compared with the video data.
[0476] Step 3:
[0477] Response Generation
[0478] The server generates a primary response using a generative AI model (e.g., GPT-4) based on the analyzed data. The input is text data and video data, and the output is the text data of the generated primary response. For example, a response is generated based on the prompt sentence, "The user is experiencing a payment error. He is in a stressful state. Please generate an appropriate customer support response."
[0479] Step 4:
[0480] Response Send
[0481] The generated response is converted from text to speech. Text-to-speech (TTS) technology (e.g., Amazon Polly) is used for the speech conversion. The response is then sent to the user through the interface system via the response sending means. The input is the text data of the generated response, and the output is audio data. The text data is converted to audio data and transmitted to the user through the interface system.
[0482] Step 5:
[0483] notification
[0484] The server notifies the mobile device of the user's information along with the generated response. Notifications are made using a real-time notification system such as Firebase Cloud Messaging. The input is the response and user information, and the output is notification data sent to the mobile device. The notification system sends notifications to the user's smartphone in real time.
[0485] Step 6:
[0486] User response reception
[0487] The mobile terminal receives the notification and provides an interface for the user to respond or select an automatic response. The input is the notification data and the output is the user's selection. The user can review the response and select either a manual or automatic response.
[0488] Step 7:
[0489] Auto-response generation
[0490] The server generates an automatic response based on the away mode set by the user. The away response is constructed by a generation AI and includes a message informing the user that they are away. The input is the away mode status, and the output is the auto-response message. An away auto-response is generated and sent to the visitor.
[0491] Step 8:
[0492] Risk Assessment
[0493] The server studies past transaction data and evaluates fraud risk. If the risk is assessed as high, it issues a warning to the user. The input is past transaction data, and the output is the fraud risk assessment result and a warning notification. Past data is analyzed by an algorithm to evaluate risk.
[0494] Step 9:
[0495] Emotion analysis
[0496] The server uses an emotion engine to analyze the user's emotional state. The input is the user's video and audio data, and the output is the emotion analysis results. For example, it can detect the user's stress or anxiety.
[0497] Step 10:
[0498] Response Adjustment
[0499] The server adjusts the response based on the emotion analysis results. If the user is feeling stressed, it generates a response with a tone that gives a sense of security. The input is the emotion analysis result, and the output is the adjusted response. A response with an appropriate tone according to the emotion is generated and sent to the user.
[0500] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0501] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0502] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0503] [Second embodiment]
[0504] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0505] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0506] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0507] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0508] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0509] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0510] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0511] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0512] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0513] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0514] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0515] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0516] The present invention is a system for automating safe and efficient visitor attendance related to an intercom system. The system is implemented using the following components:
[0517] System configuration
[0518] 1. Signal receiving means:
[0519] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[0520] 2. Data analysis methods:
[0521] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[0522] 3. Response Generation Method:
[0523] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[0524] 4. Response sending method:
[0525] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[0526] 5. Means of notification:
[0527] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[0528] 6. User response acceptance means:
[0529] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether to respond to it themselves.
[0530] 7. Automated Response Methods:
[0531] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[0532] 8. Learning and assessment tools:
[0533] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[0534] 9. Shut-out measures:
[0535] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[0536] Explanation of program processing
[0537] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[0538] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[0539] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[0540] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[0541] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[0542] This system reduces the burden on users in dealing with visitors and allows them to communicate with them safely and efficiently.
[0543] The processing flow will be explained below.
[0544] Step 1:
[0545] A visitor presses the intercom
[0546] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[0547] Step 2:
[0548] Sending data to the server
[0549] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[0550] Step 3:
[0551] First-order response by generative AI
[0552] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[0553] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[0554] The server uses the generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[0555] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[0556] Step 4:
[0557] Notification of visitor information to mobile app
[0558] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[0559] Step 5:
[0560] User response decision
[0561] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[0562] The user can then decide whether to respond based on that information.
[0563] Step 6:
[0564] User response
[0565] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[0566] The user can connect to the intercom using the app and begin talking directly with the visitor.
[0567] Step 7:
[0568] Out of Office Replies
[0569] The server will automatically generate the appropriate response if the user has configured unattended mode.
[0570] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[0571] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[0572] Step 8:
[0573] Learn visitor characteristics and identify suspicious individuals
[0574] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[0575] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[0576] If the risk score is high, the server will notify the user with a warning.
[0577] Step 9:
[0578] Shutdown execution
[0579] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[0580] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[0581] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[0582] Example 1
[0583] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0584] In today's society, where safety and convenience are essential, the automation of visitor responses is becoming increasingly important. However, with conventional intercom systems, responses are performed manually, which places a burden on busy users. Furthermore, when a suspicious person visits, an immediate and appropriate response is required, but manual responses can be risky. Therefore, there is an urgent need to develop a system that automates visitor responses over intercoms in a safe and efficient manner.
[0585] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0586] In this invention, the server includes receiving means for receiving intercom signals, analyzing means for analyzing video and audio data acquired from the intercom and identifying visitor information, generating means for generating a primary response using a generative AI model based on the analyzed data, transmitting means for sending the generated response to the visitor via the intercom, notifying means for notifying a mobile device of the visitor information and the generated response content, receiving means for accepting a response decision from the mobile device, automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, evaluation means for learning visitor characteristics and comparing them with past data to evaluate suspicious person risk, and means for notifying the mobile device of the evaluated visitor information. This enables automated visitor response, safe and efficient visitor management, and rapid response to suspicious person risks.
[0587] The "receiving means" is a device or program for receiving an intercom signal and acquiring video and audio data.
[0588] The "analysis means" is a device or program that has the function of analyzing the video and audio data acquired by the receiving means and identifying information about the visitor.
[0589] The "generation means" is a device or program that generates a primary response using a generative AI model based on the data analyzed by the analysis means.
[0590] The "transmitting means" is a device or program that transmits the response generated by the generating means to the visitor via the intercom.
[0591] The "notification means" is a device or program for notifying the mobile terminal of the visitor information and the generated response content.
[0592] The "accepting means" is a device or program that has the function of accepting a response decision from a mobile terminal.
[0593] An "automatic response means" is a device or program that has the function of generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[0594] The "assessment means" is a device or program that has the function of learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[0595] A "generative AI model" is an artificial intelligence algorithm or platform that generates appropriate responses based on analyzed data.
[0596] A "prompt" is an instruction given to a generative AI model that serves as a basis for generating a specific response.
[0597] The present invention provides an improved intercom system for automating visitor responses. The system includes a receiving unit, an analyzing unit, a generating unit, a transmitting unit, a notifying unit, a reception unit, an automatic response unit, and an evaluation unit.
[0598] Hardware and Software
[0599] 1. Receiving means
[0600] Hardware: Intercom, Server
[0601] Software: The receiving program receives signals from the intercom via the Internet and acquires video and audio data.
[0602] 2. Analysis method
[0603] Hardware: Server
[0604] Software: Uses voice recognition systems (e.g., Google Speech-to-Text) to convert voice data into text, and runs image processing algorithms (e.g., facial recognition technology) to analyze video data.
[0605] 3. Generation means
[0606] Hardware: Server
[0607] Software: Generative AI models (e.g., OpenAI GPT-3) are used to generate a first-order response based on the analyzed text and video data.
[0608] 4. Transmission Method
[0609] Hardware: Servers, intercoms
[0610] Software: Software that converts the response text into speech (e.g., a text-to-speech engine) and a transmitting program that sends it to the intercom.
[0611] 5. Means of notification
[0612] Hardware: Servers, mobile devices
[0613] Software: A program that uses the push notification function of the mobile app to notify the mobile device of visitor information and the generated response.
[0614] 6. Method of reception
[0615] Hardware: Mobile devices
[0616] Software: A mobile application that provides a user interface where users can view the response and choose whether to respond manually or via an automated response.
[0617] 7. Automated Response Methods
[0618] Hardware: Server
[0619] Software: A program that generates an automatic response message based on the absence mode set by the user, converts it into voice data, and sends it to the intercom.
[0620] 8. Evaluation Methods
[0621] Hardware: Server
[0622] Software: Machine learning algorithms that learn visitor characteristics and programs that assess risk of suspicious behavior based on historical data.
[0623] Specific processing flow
[0624] First, when a visitor presses the intercom button, the signal is sent to the server. The receiving means receives this and captures the video and audio data. Next, the server's analysis means converts the audio data into text using a voice recognition system, and analyzes the video data using facial recognition technology.
[0625] The generating means uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate primary response based on the text data and video data. This generated response is output as text data, converted into audio data by the transmitting means, and transmitted to the visitor via the intercom. Through this process, the visitor can receive the response generated by the AI.
[0626] The notification means also transmits the analysis results to the mobile device so that the user can check the response. The user can choose to respond manually using the reception means or leave it to the automatic response means. If the absence mode is set, the automatic response means automatically generates and transmits an appropriate response.
[0627] Furthermore, the evaluation means can learn visitor data, evaluate the risk of suspicious individuals, and issue a warning to the user as necessary.
[0628] Examples of specific examples and prompts
[0629] For example, if the visitor is a delivery person, the voice data of the delivery person saying "Delivery service" is sent to the server. The server uses a voice recognition system to generate the text "Delivery service," and the generative AI model then generates a response such as "Hello, you are a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom.
[0630] Examples of prompts include:
[0631] "Please tell me how you would respond if the visitor was a delivery person."
[0632] Please explain how to respond if a suspicious person visits.
[0633] This reduces the burden on users in dealing with visitors and allows them to communicate with visitors safely and efficiently.
[0634] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0635] Step 1:
[0636] The visitor presses the intercom button.
[0637] Input: Visitor Action
[0638] Output: Signal from intercom to server
[0639] Specific operation: When the intercom button is pressed, video and audio data is generated and sent to the server along with a signal.
[0640] Step 2:
[0641] The server receives the signal transmitted from the interphone via the Internet.
[0642] Input: Intercom signal, video and audio data
[0643] Output: Video and audio data stored in a database
[0644] Specific operation: The server's receiving means captures the signal and stores the video and audio data in a database.
[0645] Step 3:
[0646] The server analyzes the received voice data and converts it into text using a voice recognition system (e.g., Google Speech-to-Text).
[0647] Input: Received audio data
[0648] Output: Text data
[0649] Specific operations: The server's analysis means analyzes the voice data using a voice recognition system and generates corresponding text data.
[0650] Step 4:
[0651] The server analyzes the received video data and attempts to identify the visitor using facial recognition technology.
[0652] Input: Received video data
[0653] Output: Visitor's specific information (face image, identification information)
[0654] What it does: The server's analytics runs the video data through image processing algorithms to recognize and identify the visitor's face and stores that information.
[0655] Step 5:
[0656] The server generates a first-order response based on the analyzed text and video data using a generative AI model (e.g., OpenAI GPT-3).
[0657] Input: Text data, video data
[0658] Output: Text data of the primary responses
[0659] Specific operation: The server's generation means inputs text data and video data into the generative AI model and generates a first response such as "Hello, how can I help you?"
[0660] Step 6:
[0661] The server converts the generated primary response into voice data and transmits it to the intercom.
[0662] Input: Text data of the primary response
[0663] Output: Audio data, sent to intercom
[0664] Specific operation: The server's transmission means converts the primary response into voice data using a text-to-speech engine, and transmits the voice data to the intercom.
[0665] Step 7:
[0666] The server notifies the mobile terminal of the analysis results and the generated response.
[0667] Input: Visitor information, content of initial response
[0668] Output: Notification message to mobile device
[0669] Specific operation: The server's notification means uses a push notification service to send the visitor's video, part of the audio text, and the response content to the mobile device.
[0670] Step 8:
[0671] The mobile terminal provides an interface for the user to check the notification content and select a manual or automatic response.
[0672] Input: Notification message from the server
[0673] Output: User response selection (manual or automatic)
[0674] What it does: Displays an interface on the mobile device that allows the user to view the notification and then provide options for how to respond.
[0675] Step 9:
[0676] If the user has enabled away mode, the server automatically generates an appropriate response and sends it to the visitor.
[0677] Input: User's away mode setting
[0678] Output: Auto-answer message, sent to intercom
[0679] Specific operation: The server's automatic response means uses the generative AI model to generate a message such as "I'm not here right now," converts it into voice data, and sends it to the intercom.
[0680] Step 10:
[0681] The server learns the visitor's characteristic data and compares it with past data to assess the risk of suspicious behavior.
[0682] Input: Visitor video and audio data, past visitor data
[0683] Output: Risk assessment results, warning notifications if necessary
[0684] Specific operation: The server's evaluation means uses a machine learning algorithm to analyze visitor data, assess the risk of suspicious activity, and if the risk is high, issue a warning to the user.
[0685] This detailed processing step ensures visitor interaction is automated, secure, and efficient.
[0686] (Application example 1)
[0687] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0688] With conventional intercom systems, communication with visitors is done manually, which places a heavy burden on the person in charge and makes it difficult to respond, especially when the person in charge is not present. Furthermore, in busy environments such as logistics centers, efficient visitor response is required. Furthermore, it is difficult to assess the risk of suspicious individuals, and measures to ensure safety are insufficient. To solve these issues, a system that automates visitor response safely and efficiently is needed.
[0689] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0690] In this invention, the server includes a signal receiving means for receiving an intercom signal, a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information, a response generation means for generating a primary response using a generative AI model based on the analyzed data, a response transmission means for sending the generated response to the visitor via the intercom, a notification means for notifying a mobile device of the visitor information and the generated response, a user response receiving means for receiving a response decision from the mobile device, an automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, a learning and evaluation means for learning visitor characteristics and comparing them with past data to evaluate the risk of suspicious activity, a means for converting visitor information into text using a voice recognition system, a means for generating a primary response based on the text data using a generative AI model, a means for converting the generated text response into voice data, a communication means for notifying a smartphone or tablet of the analyzed visitor information, and a means for providing a user interface for determining a response based on the notified information. This reduces the burden on visitors and enables safe and efficient visitor response.
[0691] The "signal receiving means" is a means for receiving a signal from the interphone and processing the signal.
[0692] The "data analysis means" is a means for analyzing the video and audio data acquired from the intercom and identifying the visitor's information.
[0693] A "response generation means" is a means for generating a primary response using a generative AI model based on the analyzed data.
[0694] The "response transmitting means" is a means for transmitting the generated response to the visitor via the intercom.
[0695] The "notification means" is a means for notifying the mobile terminal of the visitor information and the generated response content.
[0696] The "user response receiving means" is a means for receiving a response decision from a mobile terminal.
[0697] The "automatic response means" is a means for generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[0698] The "learning and evaluation means" is a means for learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[0699] A "voice recognition system" is a system for converting visitor information from voice data into text.
[0700] A "generative AI model" is an artificial intelligence model for generating first-order responses based on text data.
[0701] "Speech conversion means" refers to means for converting the generated text response into voice data.
[0702] "Communication means" refers to the means for notifying analyzed visitor information to a smartphone or tablet.
[0703] A "user interface" is a means for providing an interface for determining a response based on notified information.
[0704] To put the present invention into practice, the following system and its operation will be described. This system automates visitor reception at a logistics center in a safe and efficient manner.
[0705] First, a signal receiving means is provided to receive an intercom signal. This signal receiving means serves to acquire video and audio data generated by pressing the intercom button. Next, a data analysis means analyzes the video and audio data acquired from the intercom and identifies visitor information. The identified information is processed by a response generation means that generates a primary response using a generative AI model based on the analyzed data.
[0706] The generated primary response is sent to the visitor via the intercom via the response sending means. At the same time, the visitor information and the generated response are sent to the administrator's mobile device via the notification means. The mobile device is assumed to be a smartphone or tablet.
[0707] The mobile terminal is provided with a user response reception means, and the administrator can decide whether to respond manually or continue with an automatic response based on the received notification. The response decision may also be made automatically based on the absence mode set by the user. In this case, the automatic response means generates an appropriate response and sends it to the visitor.
[0708] In addition, the server has a learning and evaluation means that learns the characteristics of visitors and compares them with past data to evaluate the risk of suspicious persons. If the risk of suspicious persons is evaluated as high, a warning is sent to the user and, if necessary, a shut-out instruction is executed. A shut-out message is sent to the visitor via the intercom.
[0709] The system also includes a means for converting visitor information into text using a speech recognition system. For example, the Google Cloud Speech-to-Text API is used to convert the visitor's speech into text. A primary response based on the text data is generated using a generative AI model (e.g., OpenAI GPT-4). The generated text response is converted into voice data using a speech conversion means (e.g., Amazon Polly).
[0710] The server also has a communication method for notifying the smartphone or tablet of the analyzed visitor information. This communication is performed using a service such as Firebase Cloud Messaging (FCM). A UI framework such as React Native is used to provide a user interface for determining a response based on the notified information.
[0711] For example, if a visitor says, "Today's delivery," this voice data is converted into text data, "Today's delivery," via the Google Cloud Speech-to-Text API. A generative AI model receives this text data and generates a response, for example, "Hello, you're the delivery person. Please wait a moment." This response is converted into voice data by Amazon Polly and conveyed to the visitor via the intercom. At the same time, this information is also notified to the administrator's smartphone.
[0712] Example prompt: "Generate an appropriate response if the visitor says, 'Delivery today.'"
[0713] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0714] Step 1:
[0715] When the interphone button is pressed, the signal receiving means receives a signal from the interphone. This signal includes video data and audio data. After receiving the signal, these data are transmitted to the server.
[0716] Input: Intercom signal
[0717] Output: Video data, audio data
[0718] Step 2:
[0719] The server analyzes the received video and audio data using a data analysis tool. The audio data is converted to text using the Google Cloud Speech-to-Text API. The video data is analyzed using a facial recognition algorithm to obtain visitor information.
[0720] Input: Video data, audio data
[0721] Output: Text data, visitor information (face recognition results)
[0722] Step 3:
[0723] The server generates a first response using a generative AI model based on the analyzed text data and visitor information. OpenAI GPT-4 is used for this. The generated first response is output as text data.
[0724] Input: Text data, visitor information
[0725] Output: Text data of the primary response
[0726] Step 4:
[0727] The server converts the generated text data of the primary response into voice data using Amazon Polly as a voice conversion means.
[0728] Input: Text data of the primary response
[0729] Output: Audio data of the first response
[0730] Step 5:
[0731] The server transmits the generated voice data of the primary response to the intercom through the response transmitting means, thereby notifying the visitor of the response.
[0732] Input: Primary response audio data
[0733] Output: Visitor receives a voice response
[0734] Step 6:
[0735] The server notifies the administrator of the analyzed visitor information and the generated primary response content via a notification means to the administrator's smartphone or tablet.
[0736] Input: Visitor information, temporary response text data
[0737] Output: Notification to the administrator's smartphone
[0738] Step 7:
[0739] When the administrator receives a notification on their smartphone or tablet, the visitor's video, audio, and initial response are displayed, allowing the user to decide whether to respond manually or select an automatic response.
[0740] Input: Notified information (visitor's video, audio text, initial response)
[0741] Output: User response prompt
[0742] Step 8:
[0743] If the user judges that the visitor is a high risk of being a suspicious person, the server uses learning and evaluation means to evaluate the risk, and if necessary, issues a warning to the administrator and executes a shut-out instruction. A shut-out message is sent to the visitor via the intercom.
[0744] Input: Information about suspicious person risk, user shut-out instructions
[0745] Output: Shutout message
[0746] Step 9:
[0747] The server analyzes and learns from visitor characteristics and past data via learning and evaluation means, and reflects this in future responses, thereby improving the accuracy of future visitor risk assessments.
[0748] Input: Visitor characteristics, historical data
[0749] Output: Learning results, updated evaluation model
[0750] In this way, visitor reception can be automated, safely, and efficiently achieved.
[0751] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0752] The present invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components.
[0753] System configuration
[0754] 1. Signal receiving means:
[0755] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[0756] 2. Data analysis methods:
[0757] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[0758] 3. Response Generation Method:
[0759] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[0760] 4. Response sending method:
[0761] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[0762] 5. Means of notification:
[0763] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[0764] 6. User response acceptance means:
[0765] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether or not to respond to it.
[0766] 7. Automated Response Methods:
[0767] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[0768] 8. Learning and assessment tools:
[0769] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[0770] 9. Shut-out measures:
[0771] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[0772] 10. Emotion Engine:
[0773] The server is equipped with an emotion engine that analyzes the user's video and audio data to recognize the user's emotions. The emotion engine analyzes the emotions (e.g., stress, anxiety, joy, surprise, etc.) that the user shows while interacting with visitors and adjusts the system's response based on that information.
[0774] 11. Emotion-based response adjustment measures:
[0775] The server adjusts the operation of the notification means and response generation means based on the user's emotions recognized by the emotion engine. For example, if the user is feeling stressed or anxious, the response generation means generates a response with a gentler tone or switches to an automatic response.
[0776] Explanation of program processing
[0777] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[0778] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[0779] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[0780] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[0781] Furthermore, the server analyzes the user's video and audio using an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed, the emotion engine detects this and the server adjusts the response generation means based on that information. As a result, the generated response is delivered in an appropriate tone according to the user's emotional state.
[0782] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[0783] This system reduces the burden on users when dealing with visitors and allows them to communicate with them safely and efficiently. In addition, the introduction of an emotion engine enables responses that take into account the user's emotional state, providing even greater peace of mind and convenience.
[0784] The processing flow will be explained below.
[0785] Step 1:
[0786] A visitor presses the intercom
[0787] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[0788] Step 2:
[0789] Sending data to the server
[0790] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[0791] Step 3:
[0792] First-order response by generative AI
[0793] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[0794] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[0795] The server uses a generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[0796] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[0797] Step 4:
[0798] Notification of visitor information to mobile app
[0799] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[0800] Step 5:
[0801] User response decision
[0802] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[0803] The user can then decide whether to respond based on that information.
[0804] Step 6:
[0805] User response
[0806] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[0807] The user can connect to the intercom using the app and begin talking directly with the visitor.
[0808] Step 7:
[0809] Out of Office Replies
[0810] The server will automatically generate the appropriate response if the user has configured unattended mode.
[0811] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[0812] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[0813] Step 8:
[0814] Learn visitor characteristics and identify suspicious individuals
[0815] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[0816] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[0817] If the risk score is high, the server will notify the user with a warning.
[0818] Step 9:
[0819] Shutdown execution
[0820] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[0821] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[0822] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[0823] Step 10:
[0824] Recognizing user emotions with an emotion engine
[0825] The server acquires the user's video and audio data from the user's mobile terminal.
[0826] The server uses an emotion engine to analyze the user's video and audio data and recognize the user's emotions.
[0827] Step 11:
[0828] Emotion-Based Response Modulation
[0829] The server adjusts the response generation means based on the user's emotions recognized by the emotion engine.
[0830] If the user shows signs of stress or anxiety, the server generates a response in a gentle tone using the response generating means, or switches to an automatic response.
[0831] Example 2
[0832] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0833] Conventional intercom systems require users to respond to visitors manually, placing a heavy burden on users. It is also difficult to respond appropriately when the resident is away or to dangerous visitors, resulting in safety and convenience issues. Furthermore, the system is unable to respond flexibly to the user's emotional state, which can be stressful.
[0834] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a signal receiving means for receiving an intercom signal; a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information; a response generation means for generating a primary response using a generative model based on the analyzed data; a response transmission means for transmitting the generated response to the visitor via the intercom; a notification means for notifying the mobile communication device of the visitor information and the generated response content; a user response receiving means for receiving a response decision from the mobile communication device; an automatic response means for generating an automatic response based on an absence mode set by the user and transmitting it to the visitor; a learning and evaluation means for learning visitor characteristics and evaluating the risk of a suspicious person by comparing them with past data; a means for analyzing the user's video and audio data and having an emotion engine that recognizes the user's emotions; and an emotion-based response adjustment means for adjusting the operation of the notification means and the response generation means based on the recognized user emotions. This automates visitor response, reducing the burden on the user and enabling appropriate responses to be taken when the user is absent or when a dangerous visitor is present. Furthermore, a system that can respond flexibly according to the user's emotional state is provided, allowing the system to be used with peace of mind.
[0835] The "signal receiving means" is a device or function for receiving a signal transmitted from the intercom.
[0836] "Data analysis means" refers to a device or function that analyzes the video and audio data acquired from the intercom and identifies visitor information.
[0837] A "response generator" is a device or function that uses a generative model based on analyzed data to generate a primary response.
[0838] The "response transmitting means" is a device or function for transmitting the generated response to the visitor via the intercom.
[0839] The "notification means" refers to a device or function for notifying the portable communication device of the visitor's information and the generated response content.
[0840] The "user response receiving means" is a device or function for receiving a response decision from the portable communication device.
[0841] The "automatic response means" is a device or function that generates an automatic response based on the absence mode set by the user and sends it to the visitor.
[0842] "Learning and evaluation means" refers to devices and functions that learn the characteristics of visitors and compare them with past data to evaluate the risk of suspicious persons.
[0843] An "emotion engine" is a device or function that analyzes a user's video and audio data and recognizes the user's emotions.
[0844] The "emotion-based response adjustment means" is a device or function for adjusting the operation of the notification means or response generation means based on the recognized user emotion.
[0845] This invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining it with an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components:
[0846] System configuration
[0847] 1. Signal receiving means:
[0848] The server receives the intercom signal and captures the video and audio data. When a visitor presses the intercom, the signal is sent to the server via the Internet. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server.
[0849] 2. Data analysis methods:
[0850] The server receives the video and audio data sent from the intercom. It converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text, and then compares it with the video data. For example, a voice saying "This is a delivery" can be converted into text like "Takhaibindesu."
[0851] 3. Response Generation Method:
[0852] The server generates an appropriate response using a generative AI model (e.g., OpenAI's GPT-4) based on the text and video data from the voice. The generated response is first output in text format and then reconverted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[0853] 4. Response sending method:
[0854] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[0855] 5. Means of notification:
[0856] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The notification includes an image of the visitor's face, a portion of the voice text, and the response. For example, the user's smartphone may receive a notification saying, "A delivery person has visited you. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[0857] 6. User response acceptance means:
[0858] The mobile device displays the received notification for the user to review. The user can then review the notification and choose to respond to it themselves, have it automatically respond, or shut it out. For example, if the user is busy, they can choose to automatically respond.
[0859] 7. Automated Response Methods:
[0860] If the user is in away mode, the server automatically generates an appropriate response to send to the intercom, such as "I'm not here right now. Please come back later."
[0861] 8. Learning and assessment tools:
[0862] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, it sends a warning to the user.
[0863] 9. Shut-out measures:
[0864] Users can select the shut-out option on their mobile app, in which case the server will generate a shut-out message and send it to the visitor over the intercom, such as "Please refrain from visiting."
[0865] 10. Emotion Engine:
[0866] The server is equipped with an emotion engine that analyzes the user's video and audio data and recognizes the user's emotions, such as stress, anxiety, joy, surprise, etc.
[0867] 11. Emotion-based response adjustment measures:
[0868] The server adjusts the content and tone of the generative AI model's responses based on the user's perceived emotions. For example, if the user is feeling stressed, the server can generate responses in a gentler tone or switch to an automatic response mode.
[0869] Examples of specific examples and prompts
[0870] For example, consider a scenario where a delivery person arrives.
[0871] 1. Signal reception: The delivery person presses the intercom, and a "ping pong" signal is sent to the server.
[0872] 2. Analysis of audio and video data: The server converts the audio "This is a delivery" into the text "Takhaibindesu" and recognizes the visitor's face.
[0873] 3. Generate a response: The generative AI model (e.g., GPT-4) generates a response such as, "Hello, you're a delivery person. Please wait a moment."
[0874] 4. Sending response: The server converts the response into voice data and sends it back to the intercom to be conveyed to the delivery person.
[0875] 5. User notification: The server notifies the user of the visitor information and the response to their mobile device. The user receives a notification that a delivery person has visited them. Response: "Hello, it's you, a delivery person. Please wait a moment."
[0876] 6. Accepting user response: The user opens the app and selects how to respond. Since they are busy, they select automatic response.
[0877] 7. Execute Auto-Reply (Away Mode): If the user has Away Mode enabled, a message will be generated saying "I'm out of the office right now. Please come back later."
[0878] 8. Learning and risk assessment: The server learns from the video and audio of the delivery person and performs risk assessment. It determines that the person is not suspicious.
[0879] 9. Execute Shutout: Send a shutout message depending on conditions such as absence or risk of suspicious person. Not necessary this time.
[0880] 10. Analysis by emotion engine: The emotions (e.g., anxiety) that users feel when interacting with visitors are analyzed.
[0881] 11. Adjusting responses based on emotion: If the server is stressed, it will generate a gentler response such as "Please wait a moment."
[0882] For example, the prompt to input to a generative AI model might look like this:
[0883] Caller's voice: "This is a courier."
[0884] Prompt to generative AI: "The visitor says 'This is a courier.' Please generate an appropriate first-level response."
[0885] This allows the system to provide highly efficient visitor support.
[0886] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0887] Step 1: Signal reception
[0888] When a visitor presses the intercom, the server receives the signal via the Internet. This signal contains the visitor's video and audio data. The input is the signal when the visitor presses the intercom, and the output is the video and audio data stored on the server. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server, and the server captures the video and audio at that moment.
[0889] Step 2: Data analysis
[0890] The server analyzes the received video and audio data. First, it converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text. Next, it analyzes the video data and recognizes the visitor's face. This converts the input audio data into text format and also generates facial recognition data. For example, the audio saying "This is a delivery person" is converted into the text "Takhaibindesu" and the visitor's face is identified.
[0891] Step 3: Response Generation
[0892] The server generates an appropriate first-order response using a generative AI model (e.g., OpenAI's GPT-4) based on the text data and facial recognition data. The input is the text-converted voice data and facial recognition data, and the output is the response text. This response text is then converted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[0893] Step 4: Send response
[0894] The server sends the generated voice data to the intercom, and the visitor can receive the generated response. The input is the voice data generated by the server, and the output is the voice response transmitted to the visitor through the intercom. For example, the voice transmitted to the delivery person might say, "Hello, you are a delivery person. Please wait a moment."
[0895] Step 5: Notification
[0896] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The input is the visitor information (face image, part of the voice text) and the generated response, and the output is a notification message. A notification containing the visitor's face image and text is displayed on the user's mobile device. For example, the notification may read, "A delivery person has visited. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[0897] Step 6: Accept user response
[0898] The device displays the received notification for the user to review. The user can then review it and choose to respond manually, have it automatically respond, or shut it out. The input is the notification message from the server, and the output is the user's response decision. For example, if the user is busy and chooses to have it automatically respond, they can select it through the app.
[0899] Step 7: Auto-responders
[0900] If the user is in away mode, the server automatically generates an appropriate response and sends it to the intercom. The input is the away mode setting, and the output is an automatic response message. For example, a message like "I'm not available right now. Please come back later" is generated.
[0901] Step 8: Learning and risk assessment
[0902] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. The input is past visitor data and current visitor data, and the output is the suspicious person risk assessment result. If the server determines there is a high risk, it sends a warning to the user.
[0903] Step 9: Shut Out
[0904] The user can select the shut-out option on the mobile app. In this case, the server generates a shut-out message and transmits it to the visitor through the intercom. The input is the user's shut-out instruction, and the output is the shut-out message. For example, a message such as "Please refrain from visiting" is generated.
[0905] Step 10: Emotion Engine Analysis
[0906] The server analyzes the user's video and audio data and recognizes the user's emotions. The input is the user's video and audio data, and the output is the recognized emotional information. For example, it may be recognized from the user's video that they are feeling stressed.
[0907] Step 11: Adjust your response based on your emotions
[0908] The server adjusts the behavior of the notification and response generation means based on the recognized user emotion. The input is the recognized emotion information, and the output is a response adjusted according to the emotion. For example, if the user is feeling stressed, the generative AI model can generate a gentler tone of response or switch to automatic response mode.
[0909] (Application example 2)
[0910] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0911] With conventional intercom systems, responding to visitors is done manually, which is not only time-consuming for users but also poses issues in terms of safety and efficiency. Even with electronic payment services, when a problem occurs during a transaction, the response is often delayed, causing frustration for users. Furthermore, insufficient evaluation of fraud risk leaves users with security concerns.
[0912] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means. This automates visitor response and transaction management, enabling safe and efficient responses while taking into consideration the user's emotional state.
[0913] The "signal receiving means" is a function that allows the server to receive signals from the interface system.
[0914] The "data analysis means" is a function for analyzing the video and audio data acquired from the interface system to identify user information.
[0915] The "response generation means" is a function for generating a primary response using a generative AI model based on the analyzed data.
[0916] The "response sending means" is a function for sending the generated response to the user through the interface system.
[0917] The "notification means" is a function for notifying the mobile terminal of the user information and the generated response content.
[0918] The "user response receiving means" is a function for receiving a response decision from a mobile terminal.
[0919] The "automatic response means" is a function for generating an automatic response based on the absence mode set by the user and sending it to the user.
[0920] The "learning and evaluation means" is a function for learning from past transaction data and evaluating fraud risk.
[0921] The "emotion analysis means" is a function for analyzing the emotional state of the user.
[0922] The "response adjustment means" is a function for adjusting the response based on the analysis results of the emotion analysis means.
[0923] The present invention provides a system for automating visitor response and transaction management, enabling safe and efficient response while taking into consideration the emotional state of the user. The system includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means.
[0924] System configuration and operation
[0925] 1. Signal receiving means:
[0926] The server receives a signal from an interface system (e.g., an electronic payment terminal) to obtain transaction data.
[0927] 2. Data analysis methods:
[0928] The server analyzes the video and audio data acquired from the interface system to identify the user's information. Specifically, it converts the audio data into text using a speech recognition system (e.g., Google Cloud Speech-to-Text) and analyzes it.
[0929] 3. Response Generation Method:
[0930] A generative AI model (e.g., GPT-4) is used to generate a first-order response based on the analyzed data. Example prompts for response generation are as follows:
[0931] "User is experiencing payment error. Stressed. Please generate an appropriate customer support response."
[0932] 4. Response sending method:
[0933] The generated response is converted into text or speech and sent to the user through the interface system, using text-to-speech (TTS) technology (e.g., Amazon Polly) to reconvert the speech.
[0934] 5. Means of notification:
[0935] The server then sends the user information and the generated response to the mobile device (smartphone) using a real-time notification system such as Firebase Cloud Messaging.
[0936] 6. User response acceptance means:
[0937] The mobile terminal provides an interface for receiving the notification and accepting the user's response. The user can respond manually or select an automatic response.
[0938] 7. Automated Response Methods:
[0939] Based on the out-of-office mode set by the user, an automated response is generated and sent to the user. The out-of-office response is constructed by the generative AI and includes a message informing the user that they are out of the office.
[0940] 8. Learning and assessment tools:
[0941] The server has an algorithm that learns from past transaction data and evaluates fraud risk. If the fraud risk is assessed as high, it will send a warning to the user.
[0942] 9. Emotion analysis means:
[0943] The server uses an emotion engine (e.g., IBM Watson or Microsoft Azure Emotion API) to analyze the user's emotional state, detecting emotions such as stress, anxiety, and joy.
[0944] 10. Response Adjustment Measures:
[0945] The server adjusts the response based on the analysis results of the emotion analysis means, and if the user is feeling stressed, it generates a response with a more reassuring tone.
[0946] Examples:
[0947] For example, consider a scenario in which a user encounters a payment error while using an electronic payment service and seeks support. A signal is sent from the interface system to the server, which converts the voice data into text. The server generates an appropriate response, converts it into voice, and sends it to the user. At the same time, the mobile device is notified of the details of the problem and the generated response. Furthermore, the server uses an emotion engine to analyze the user's emotional state and adjust the response accordingly. An example of a specific prompt would be, "The user is experiencing a payment error. He is in a stressful state. Please generate an appropriate customer support response."
[0948] This system reduces the burden on users in dealing with visitors and managing transactions, allowing them to receive safe and efficient services. In addition, the introduction of an emotion engine makes it possible to respond according to the user's emotional state, providing even greater peace of mind and convenience.
[0949] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0950] Step 1:
[0951] Signal Reception
[0952] The server receives signals from the interface system (e.g., electronic payment terminal). These signals include visitor information and transaction data. The input is the signal from the interface system, and the output is the received signal data. The signal sent from the interface system is received by the server's API.
[0953] Step 2:
[0954] Data analysis
[0955] The server analyzes the received video and audio data. The audio data is first converted into text using a speech recognition system (e.g., Google Cloud Speech-to-Text). Analysis is performed based on this text data and video data. The input is the received signal data, and the output is text data and analyzed video data. The audio data is converted into text format and then compared with the video data.
[0956] Step 3:
[0957] Response Generation
[0958] The server generates a primary response using a generative AI model (e.g., GPT-4) based on the analyzed data. The input is text data and video data, and the output is the text data of the generated primary response. For example, a response is generated based on the prompt sentence, "The user is experiencing a payment error. He is in a stressful state. Please generate an appropriate customer support response."
[0959] Step 4:
[0960] Response Send
[0961] The generated response is converted from text to speech. Text-to-speech (TTS) technology (e.g., Amazon Polly) is used for the speech conversion. The response is then sent to the user through the interface system via the response sending means. The input is the text data of the generated response, and the output is audio data. The text data is converted to audio data and transmitted to the user through the interface system.
[0962] Step 5:
[0963] notification
[0964] The server notifies the mobile device of the user's information along with the generated response. Notifications are made using a real-time notification system such as Firebase Cloud Messaging. The input is the response and user information, and the output is notification data sent to the mobile device. The notification system sends notifications to the user's smartphone in real time.
[0965] Step 6:
[0966] User response reception
[0967] The mobile terminal receives the notification and provides an interface for the user to respond or select an automatic response. The input is the notification data and the output is the user's selection. The user can review the response and select either a manual or automatic response.
[0968] Step 7:
[0969] Auto-response generation
[0970] The server generates an automatic response based on the away mode set by the user. The away response is constructed by a generation AI and includes a message informing the user that they are away. The input is the away mode status, and the output is the auto-response message. An away auto-response is generated and sent to the visitor.
[0971] Step 8:
[0972] Risk Assessment
[0973] The server studies past transaction data and evaluates fraud risk. If the risk is assessed as high, it issues a warning to the user. The input is past transaction data, and the output is the fraud risk assessment result and a warning notification. Past data is analyzed by an algorithm to evaluate risk.
[0974] Step 9:
[0975] Emotion analysis
[0976] The server uses an emotion engine to analyze the user's emotional state. The input is the user's video and audio data, and the output is the emotion analysis results. For example, it can detect the user's stress or anxiety.
[0977] Step 10:
[0978] Response Adjustment
[0979] The server adjusts the response based on the emotion analysis results. If the user is feeling stressed, it generates a response with a tone that gives a sense of security. The input is the emotion analysis result, and the output is the adjusted response. A response with an appropriate tone according to the emotion is generated and sent to the user.
[0980] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0981] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0982] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0983] [Third embodiment]
[0984] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0985] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0986] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0987] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0988] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0989] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0990] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0991] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0992] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0993] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0994] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0995] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0996] The present invention is a system for automating safe and efficient visitor attendance related to an intercom system. The system is implemented using the following components:
[0997] System configuration
[0998] 1. Signal receiving means:
[0999] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[1000] 2. Data analysis methods:
[1001] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[1002] 3. Response Generation Method:
[1003] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[1004] 4. Response sending method:
[1005] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[1006] 5. Means of notification:
[1007] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[1008] 6. User response acceptance means:
[1009] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether to respond to it themselves.
[1010] 7. Automated Response Methods:
[1011] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[1012] 8. Learning and assessment tools:
[1013] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[1014] 9. Shut-out measures:
[1015] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[1016] Explanation of program processing
[1017] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[1018] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[1019] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[1020] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[1021] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[1022] This system reduces the burden on users in dealing with visitors and allows them to communicate with them safely and efficiently.
[1023] The processing flow will be explained below.
[1024] Step 1:
[1025] A visitor presses the intercom
[1026] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[1027] Step 2:
[1028] Sending data to the server
[1029] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[1030] Step 3:
[1031] First-order response by generative AI
[1032] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[1033] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[1034] The server uses the generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[1035] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[1036] Step 4:
[1037] Notification of visitor information to mobile app
[1038] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[1039] Step 5:
[1040] User response decision
[1041] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[1042] The user can then decide whether to respond based on that information.
[1043] Step 6:
[1044] User response
[1045] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[1046] The user can connect to the intercom using the app and begin talking directly with the visitor.
[1047] Step 7:
[1048] Out of Office Replies
[1049] The server will automatically generate the appropriate response if the user has configured unattended mode.
[1050] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[1051] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[1052] Step 8:
[1053] Learn visitor characteristics and identify suspicious individuals
[1054] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[1055] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[1056] If the risk score is high, the server will notify the user with a warning.
[1057] Step 9:
[1058] Shutdown execution
[1059] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[1060] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[1061] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[1062] Example 1
[1063] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1064] In today's society, where safety and convenience are essential, the automation of visitor responses is becoming increasingly important. However, with conventional intercom systems, responses are performed manually, which places a burden on busy users. Furthermore, when a suspicious person visits, an immediate and appropriate response is required, but manual responses can be risky. Therefore, there is an urgent need to develop a system that automates visitor responses over intercoms in a safe and efficient manner.
[1065] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1066] In this invention, the server includes receiving means for receiving intercom signals, analyzing means for analyzing video and audio data acquired from the intercom and identifying visitor information, generating means for generating a primary response using a generative AI model based on the analyzed data, transmitting means for sending the generated response to the visitor via the intercom, notifying means for notifying a mobile device of the visitor information and the generated response content, receiving means for accepting a response decision from the mobile device, automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, evaluation means for learning visitor characteristics and comparing them with past data to evaluate suspicious person risk, and means for notifying the mobile device of the evaluated visitor information. This enables automated visitor response, safe and efficient visitor management, and rapid response to suspicious person risks.
[1067] The "receiving means" is a device or program for receiving an intercom signal and acquiring video and audio data.
[1068] The "analysis means" is a device or program that has the function of analyzing the video and audio data acquired by the receiving means and identifying information about the visitor.
[1069] The "generation means" is a device or program that generates a primary response using a generative AI model based on the data analyzed by the analysis means.
[1070] The "transmitting means" is a device or program that transmits the response generated by the generating means to the visitor via the intercom.
[1071] The "notification means" is a device or program for notifying the mobile terminal of the visitor information and the generated response content.
[1072] The "accepting means" is a device or program that has the function of accepting a response decision from a mobile terminal.
[1073] An "automatic response means" is a device or program that has the function of generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[1074] The "assessment means" is a device or program that has the function of learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[1075] A "generative AI model" is an artificial intelligence algorithm or platform that generates appropriate responses based on analyzed data.
[1076] A "prompt" is an instruction given to a generative AI model that serves as a basis for generating a specific response.
[1077] The present invention provides an improved intercom system for automating visitor responses. The system includes a receiving unit, an analyzing unit, a generating unit, a transmitting unit, a notifying unit, a reception unit, an automatic response unit, and an evaluation unit.
[1078] Hardware and Software
[1079] 1. Receiving means
[1080] Hardware: Intercom, Server
[1081] Software: The receiving program receives signals from the intercom via the Internet and acquires video and audio data.
[1082] 2. Analysis method
[1083] Hardware: Server
[1084] Software: Uses voice recognition systems (e.g., Google Speech-to-Text) to convert voice data into text, and runs image processing algorithms (e.g., facial recognition technology) to analyze video data.
[1085] 3. Generation means
[1086] Hardware: Server
[1087] Software: Generative AI models (e.g., OpenAI GPT-3) are used to generate a first-order response based on the analyzed text and video data.
[1088] 4. Transmission Method
[1089] Hardware: Servers, intercoms
[1090] Software: Software that converts the response text into speech (e.g., a text-to-speech engine) and a transmitting program that sends it to the intercom.
[1091] 5. Means of notification
[1092] Hardware: Servers, mobile devices
[1093] Software: A program that uses the push notification function of the mobile app to notify the mobile device of visitor information and the generated response.
[1094] 6. Method of reception
[1095] Hardware: Mobile devices
[1096] Software: A mobile application that provides a user interface where users can view the response and choose whether to respond manually or via an automated response.
[1097] 7. Automated Response Methods
[1098] Hardware: Server
[1099] Software: A program that generates an automatic response message based on the absence mode set by the user, converts it into voice data, and sends it to the intercom.
[1100] 8. Evaluation Methods
[1101] Hardware: Server
[1102] Software: Machine learning algorithms that learn visitor characteristics and programs that assess risk of suspicious behavior based on historical data.
[1103] Specific processing flow
[1104] First, when a visitor presses the intercom button, the signal is sent to the server. The receiving means receives this and captures the video and audio data. Next, the server's analysis means converts the audio data into text using a voice recognition system, and analyzes the video data using facial recognition technology.
[1105] The generating means uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate primary response based on the text data and video data. This generated response is output as text data, converted into audio data by the transmitting means, and transmitted to the visitor via the intercom. Through this process, the visitor can receive the response generated by the AI.
[1106] The notification means also transmits the analysis results to the mobile device so that the user can check the response. The user can choose to respond manually using the reception means or leave it to the automatic response means. If the absence mode is set, the automatic response means automatically generates and transmits an appropriate response.
[1107] Furthermore, the evaluation means can learn visitor data, evaluate the risk of suspicious individuals, and issue a warning to the user as necessary.
[1108] Examples of specific examples and prompts
[1109] For example, if the visitor is a delivery person, the voice data of the delivery person saying "Delivery service" is sent to the server. The server uses a voice recognition system to generate the text "Delivery service," and the generative AI model then generates a response such as "Hello, you are a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom.
[1110] Examples of prompts include:
[1111] "Please tell me how you would respond if the visitor was a delivery person."
[1112] Please explain how to respond if a suspicious person visits.
[1113] This reduces the burden on users in dealing with visitors and allows them to communicate with visitors safely and efficiently.
[1114] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1115] Step 1:
[1116] The visitor presses the intercom button.
[1117] Input: Visitor Action
[1118] Output: Signal from intercom to server
[1119] Specific operation: When the intercom button is pressed, video and audio data is generated and sent to the server along with a signal.
[1120] Step 2:
[1121] The server receives the signal transmitted from the interphone via the Internet.
[1122] Input: Intercom signal, video and audio data
[1123] Output: Video and audio data stored in a database
[1124] Specific operation: The server's receiving means captures the signal and stores the video and audio data in a database.
[1125] Step 3:
[1126] The server analyzes the received voice data and converts it into text using a voice recognition system (e.g., Google Speech-to-Text).
[1127] Input: Received audio data
[1128] Output: Text data
[1129] Specific operations: The server's analysis means analyzes the voice data using a voice recognition system and generates corresponding text data.
[1130] Step 4:
[1131] The server analyzes the received video data and attempts to identify the visitor using facial recognition technology.
[1132] Input: Received video data
[1133] Output: Visitor's specific information (face image, identification information)
[1134] What it does: The server's analytics runs the video data through image processing algorithms to recognize and identify the visitor's face and stores that information.
[1135] Step 5:
[1136] The server generates a first-order response based on the analyzed text and video data using a generative AI model (e.g., OpenAI GPT-3).
[1137] Input: Text data, video data
[1138] Output: Text data of the primary responses
[1139] Specific operation: The server's generation means inputs text data and video data into the generative AI model and generates a first response such as "Hello, how can I help you?"
[1140] Step 6:
[1141] The server converts the generated primary response into voice data and transmits it to the intercom.
[1142] Input: Text data of the primary response
[1143] Output: Audio data, sent to intercom
[1144] Specific operation: The server's transmission means converts the primary response into voice data using a text-to-speech engine, and transmits the voice data to the intercom.
[1145] Step 7:
[1146] The server notifies the mobile terminal of the analysis results and the generated response.
[1147] Input: Visitor information, content of initial response
[1148] Output: Notification message to mobile device
[1149] Specific operation: The server's notification means uses a push notification service to send the visitor's video, part of the audio text, and the response content to the mobile device.
[1150] Step 8:
[1151] The mobile terminal provides an interface for the user to check the notification content and select a manual or automatic response.
[1152] Input: Notification message from the server
[1153] Output: User response selection (manual or automatic)
[1154] What it does: Displays an interface on the mobile device that allows the user to view the notification and then provide options for how to respond.
[1155] Step 9:
[1156] If the user has enabled away mode, the server automatically generates an appropriate response and sends it to the visitor.
[1157] Input: User's away mode setting
[1158] Output: Auto-answer message, sent to intercom
[1159] Specific operation: The server's automatic response means uses the generative AI model to generate a message such as "I'm not here right now," converts it into voice data, and sends it to the intercom.
[1160] Step 10:
[1161] The server learns the visitor's characteristic data and compares it with past data to assess the risk of suspicious behavior.
[1162] Input: Visitor video and audio data, past visitor data
[1163] Output: Risk assessment results, warning notifications if necessary
[1164] Specific operation: The server's evaluation means uses a machine learning algorithm to analyze visitor data, assess the risk of suspicious activity, and if the risk is high, issue a warning to the user.
[1165] This detailed processing step ensures visitor interaction is automated, secure, and efficient.
[1166] (Application example 1)
[1167] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1168] With conventional intercom systems, communication with visitors is done manually, which places a heavy burden on the person in charge and makes it difficult to respond, especially when the person in charge is not present. Furthermore, in busy environments such as logistics centers, efficient visitor response is required. Furthermore, it is difficult to assess the risk of suspicious individuals, and measures to ensure safety are insufficient. To solve these issues, a system that automates visitor response safely and efficiently is needed.
[1169] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1170] In this invention, the server includes a signal receiving means for receiving an intercom signal, a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information, a response generation means for generating a primary response using a generative AI model based on the analyzed data, a response transmission means for sending the generated response to the visitor via the intercom, a notification means for notifying a mobile device of the visitor information and the generated response, a user response receiving means for receiving a response decision from the mobile device, an automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, a learning and evaluation means for learning visitor characteristics and comparing them with past data to evaluate the risk of suspicious activity, a means for converting visitor information into text using a voice recognition system, a means for generating a primary response based on the text data using a generative AI model, a means for converting the generated text response into voice data, a communication means for notifying a smartphone or tablet of the analyzed visitor information, and a means for providing a user interface for determining a response based on the notified information. This reduces the burden on visitors and enables safe and efficient visitor response.
[1171] The "signal receiving means" is a means for receiving a signal from the interphone and processing the signal.
[1172] The "data analysis means" is a means for analyzing the video and audio data acquired from the intercom and identifying the visitor's information.
[1173] A "response generation means" is a means for generating a primary response using a generative AI model based on the analyzed data.
[1174] The "response transmitting means" is a means for transmitting the generated response to the visitor via the intercom.
[1175] The "notification means" is a means for notifying the mobile terminal of the visitor information and the generated response content.
[1176] The "user response receiving means" is a means for receiving a response decision from a mobile terminal.
[1177] The "automatic response means" is a means for generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[1178] The "learning and evaluation means" is a means for learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[1179] A "voice recognition system" is a system for converting visitor information from voice data into text.
[1180] A "generative AI model" is an artificial intelligence model for generating first-order responses based on text data.
[1181] "Speech conversion means" refers to means for converting the generated text response into voice data.
[1182] "Communication means" refers to the means for notifying analyzed visitor information to a smartphone or tablet.
[1183] A "user interface" is a means for providing an interface for determining a response based on notified information.
[1184] To put the present invention into practice, the following system and its operation will be described. This system automates visitor reception at a logistics center in a safe and efficient manner.
[1185] First, a signal receiving means is provided to receive an intercom signal. This signal receiving means serves to acquire video and audio data generated by pressing the intercom button. Next, a data analysis means analyzes the video and audio data acquired from the intercom and identifies visitor information. The identified information is processed by a response generation means that generates a primary response using a generative AI model based on the analyzed data.
[1186] The generated primary response is sent to the visitor via the intercom via the response sending means. At the same time, the visitor information and the generated response are sent to the administrator's mobile device via the notification means. The mobile device is assumed to be a smartphone or tablet.
[1187] The mobile terminal is provided with a user response reception means, and the administrator can decide whether to respond manually or continue with an automatic response based on the received notification. The response decision may also be made automatically based on the absence mode set by the user. In this case, the automatic response means generates an appropriate response and sends it to the visitor.
[1188] In addition, the server has a learning and evaluation means that learns the characteristics of visitors and compares them with past data to evaluate the risk of suspicious persons. If the risk of suspicious persons is evaluated as high, a warning is sent to the user and, if necessary, a shut-out instruction is executed. A shut-out message is sent to the visitor via the intercom.
[1189] The system also includes a means for converting visitor information into text using a speech recognition system. For example, the Google Cloud Speech-to-Text API is used to convert the visitor's speech into text. A primary response based on the text data is generated using a generative AI model (e.g., OpenAI GPT-4). The generated text response is converted into voice data using a speech conversion means (e.g., Amazon Polly).
[1190] The server also has a communication method for notifying the smartphone or tablet of the analyzed visitor information. This communication is performed using a service such as Firebase Cloud Messaging (FCM). A UI framework such as React Native is used to provide a user interface for determining a response based on the notified information.
[1191] For example, if a visitor says, "Today's delivery," this voice data is converted into text data, "Today's delivery," via the Google Cloud Speech-to-Text API. A generative AI model receives this text data and generates a response, for example, "Hello, you're the delivery person. Please wait a moment." This response is converted into voice data by Amazon Polly and conveyed to the visitor via the intercom. At the same time, this information is also notified to the administrator's smartphone.
[1192] Example prompt: "Generate an appropriate response if the visitor says, 'Delivery today.'"
[1193] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1194] Step 1:
[1195] When the interphone button is pressed, the signal receiving means receives a signal from the interphone. This signal includes video data and audio data. After receiving the signal, these data are transmitted to the server.
[1196] Input: Intercom signal
[1197] Output: Video data, audio data
[1198] Step 2:
[1199] The server analyzes the received video and audio data using a data analysis tool. The audio data is converted to text using the Google Cloud Speech-to-Text API. The video data is analyzed using a facial recognition algorithm to obtain visitor information.
[1200] Input: Video data, audio data
[1201] Output: Text data, visitor information (face recognition results)
[1202] Step 3:
[1203] The server generates a first response using a generative AI model based on the analyzed text data and visitor information. OpenAI GPT-4 is used for this. The generated first response is output as text data.
[1204] Input: Text data, visitor information
[1205] Output: Text data of the primary response
[1206] Step 4:
[1207] The server converts the generated text data of the primary response into voice data using Amazon Polly as a voice conversion means.
[1208] Input: Text data of the primary response
[1209] Output: Audio data of the first response
[1210] Step 5:
[1211] The server transmits the generated voice data of the primary response to the intercom through the response transmitting means, thereby notifying the visitor of the response.
[1212] Input: Primary response audio data
[1213] Output: Visitor receives a voice response
[1214] Step 6:
[1215] The server notifies the administrator of the analyzed visitor information and the generated primary response content via a notification means to the administrator's smartphone or tablet.
[1216] Input: Visitor information, temporary response text data
[1217] Output: Notification to the administrator's smartphone
[1218] Step 7:
[1219] When the administrator receives a notification on their smartphone or tablet, the visitor's video, audio, and initial response are displayed, allowing the user to decide whether to respond manually or select an automatic response.
[1220] Input: Notified information (visitor's video, audio text, initial response)
[1221] Output: User response prompt
[1222] Step 8:
[1223] If the user judges that the visitor is a high risk of being a suspicious person, the server uses learning and evaluation means to evaluate the risk, and if necessary, issues a warning to the administrator and executes a shut-out instruction. A shut-out message is sent to the visitor via the intercom.
[1224] Input: Information about suspicious person risk, user shut-out instructions
[1225] Output: Shutout message
[1226] Step 9:
[1227] The server analyzes and learns from visitor characteristics and past data via learning and evaluation means, and reflects this in future responses, thereby improving the accuracy of future visitor risk assessments.
[1228] Input: Visitor characteristics, historical data
[1229] Output: Learning results, updated evaluation model
[1230] In this way, visitor reception can be automated, safely, and efficiently achieved.
[1231] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1232] The present invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components.
[1233] System configuration
[1234] 1. Signal receiving means:
[1235] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[1236] 2. Data analysis methods:
[1237] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[1238] 3. Response Generation Method:
[1239] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[1240] 4. Response sending method:
[1241] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[1242] 5. Means of notification:
[1243] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[1244] 6. User response acceptance means:
[1245] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether or not to respond to it.
[1246] 7. Automated Response Methods:
[1247] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[1248] 8. Learning and assessment tools:
[1249] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[1250] 9. Shut-out measures:
[1251] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[1252] 10. Emotion Engine:
[1253] The server is equipped with an emotion engine that analyzes the user's video and audio data to recognize the user's emotions. The emotion engine analyzes the emotions (e.g., stress, anxiety, joy, surprise, etc.) that the user shows while interacting with visitors and adjusts the system's response based on that information.
[1254] 11. Emotion-based response adjustment measures:
[1255] The server adjusts the operation of the notification means and response generation means based on the user's emotions recognized by the emotion engine. For example, if the user is feeling stressed or anxious, the response generation means generates a response with a gentler tone or switches to an automatic response.
[1256] Explanation of program processing
[1257] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[1258] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[1259] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[1260] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[1261] Furthermore, the server analyzes the user's video and audio using an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed, the emotion engine detects this and the server adjusts the response generation means based on that information. As a result, the generated response is delivered in an appropriate tone according to the user's emotional state.
[1262] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[1263] This system reduces the burden on users when dealing with visitors and allows them to communicate with them safely and efficiently. In addition, the introduction of an emotion engine enables responses that take into account the user's emotional state, providing even greater peace of mind and convenience.
[1264] The processing flow will be explained below.
[1265] Step 1:
[1266] A visitor presses the intercom
[1267] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[1268] Step 2:
[1269] Sending data to the server
[1270] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[1271] Step 3:
[1272] First-order response by generative AI
[1273] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[1274] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[1275] The server uses a generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[1276] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[1277] Step 4:
[1278] Notification of visitor information to mobile app
[1279] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[1280] Step 5:
[1281] User response decision
[1282] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[1283] The user can then decide whether to respond based on that information.
[1284] Step 6:
[1285] User response
[1286] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[1287] The user can connect to the intercom using the app and begin talking directly with the visitor.
[1288] Step 7:
[1289] Out of Office Replies
[1290] The server will automatically generate the appropriate response if the user has configured unattended mode.
[1291] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[1292] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[1293] Step 8:
[1294] Learn visitor characteristics and identify suspicious individuals
[1295] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[1296] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[1297] If the risk score is high, the server will notify the user with a warning.
[1298] Step 9:
[1299] Shutdown execution
[1300] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[1301] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[1302] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[1303] Step 10:
[1304] Recognizing user emotions with an emotion engine
[1305] The server acquires the user's video and audio data from the user's mobile terminal.
[1306] The server uses an emotion engine to analyze the user's video and audio data and recognize the user's emotions.
[1307] Step 11:
[1308] Emotion-Based Response Modulation
[1309] The server adjusts the response generation means based on the user's emotions recognized by the emotion engine.
[1310] If the user shows signs of stress or anxiety, the server generates a response in a gentle tone using the response generating means, or switches to an automatic response.
[1311] Example 2
[1312] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1313] Conventional intercom systems require users to respond to visitors manually, placing a heavy burden on users. It is also difficult to respond appropriately when the resident is away or to dangerous visitors, resulting in safety and convenience issues. Furthermore, the system is unable to respond flexibly to the user's emotional state, which can be stressful.
[1314] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a signal receiving means for receiving an intercom signal; a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information; a response generation means for generating a primary response using a generative model based on the analyzed data; a response transmission means for transmitting the generated response to the visitor via the intercom; a notification means for notifying the mobile communication device of the visitor information and the generated response content; a user response receiving means for receiving a response decision from the mobile communication device; an automatic response means for generating an automatic response based on an absence mode set by the user and transmitting it to the visitor; a learning and evaluation means for learning visitor characteristics and evaluating the risk of a suspicious person by comparing them with past data; a means for analyzing the user's video and audio data and having an emotion engine that recognizes the user's emotions; and an emotion-based response adjustment means for adjusting the operation of the notification means and the response generation means based on the recognized user emotions. This automates visitor response, reducing the burden on the user and enabling appropriate responses to be taken when the user is absent or when a dangerous visitor is present. Furthermore, a system that can respond flexibly according to the user's emotional state is provided, allowing the system to be used with peace of mind.
[1315] The "signal receiving means" is a device or function for receiving a signal transmitted from the intercom.
[1316] "Data analysis means" refers to a device or function that analyzes the video and audio data acquired from the intercom and identifies visitor information.
[1317] A "response generator" is a device or function that uses a generative model based on analyzed data to generate a primary response.
[1318] The "response transmitting means" is a device or function for transmitting the generated response to the visitor via the intercom.
[1319] The "notification means" refers to a device or function for notifying the portable communication device of the visitor's information and the generated response content.
[1320] The "user response receiving means" is a device or function for receiving a response decision from the portable communication device.
[1321] The "automatic response means" is a device or function that generates an automatic response based on the absence mode set by the user and sends it to the visitor.
[1322] "Learning and evaluation means" refers to devices and functions that learn the characteristics of visitors and compare them with past data to evaluate the risk of suspicious persons.
[1323] An "emotion engine" is a device or function that analyzes a user's video and audio data and recognizes the user's emotions.
[1324] The "emotion-based response adjustment means" is a device or function for adjusting the operation of the notification means or response generation means based on the recognized user emotion.
[1325] This invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining it with an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components:
[1326] System configuration
[1327] 1. Signal receiving means:
[1328] The server receives the intercom signal and captures the video and audio data. When a visitor presses the intercom, the signal is sent to the server via the Internet. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server.
[1329] 2. Data analysis methods:
[1330] The server receives the video and audio data sent from the intercom. It converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text, and then compares it with the video data. For example, a voice saying "This is a delivery" can be converted into text like "Takhaibindesu."
[1331] 3. Response Generation Method:
[1332] The server generates an appropriate response using a generative AI model (e.g., OpenAI's GPT-4) based on the text and video data from the voice. The generated response is first output in text format and then reconverted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[1333] 4. Response sending method:
[1334] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[1335] 5. Means of notification:
[1336] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The notification includes an image of the visitor's face, a portion of the voice text, and the response. For example, the user's smartphone may receive a notification saying, "A delivery person has visited you. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[1337] 6. User response acceptance means:
[1338] The mobile device displays the received notification for the user to review. The user can then review the notification and choose to respond to it themselves, have it automatically respond, or shut it out. For example, if the user is busy, they can choose to automatically respond.
[1339] 7. Automated Response Methods:
[1340] If the user is in away mode, the server automatically generates an appropriate response to send to the intercom, such as "I'm not here right now. Please come back later."
[1341] 8. Learning and assessment tools:
[1342] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, it sends a warning to the user.
[1343] 9. Shut-out measures:
[1344] Users can select the shut-out option on their mobile app, in which case the server will generate a shut-out message and send it to the visitor over the intercom, such as "Please refrain from visiting."
[1345] 10. Emotion Engine:
[1346] The server is equipped with an emotion engine that analyzes the user's video and audio data and recognizes the user's emotions, such as stress, anxiety, joy, surprise, etc.
[1347] 11. Emotion-based response adjustment measures:
[1348] The server adjusts the content and tone of the generative AI model's responses based on the user's perceived emotions. For example, if the user is feeling stressed, the server can generate responses in a gentler tone or switch to an automatic response mode.
[1349] Examples of specific examples and prompts
[1350] For example, consider a scenario where a delivery person arrives.
[1351] 1. Signal reception: The delivery person presses the intercom, and a "ping pong" signal is sent to the server.
[1352] 2. Analysis of audio and video data: The server converts the audio "This is a delivery" into the text "Takhaibindesu" and recognizes the visitor's face.
[1353] 3. Generate a response: The generative AI model (e.g., GPT-4) generates a response such as, "Hello, you're a delivery person. Please wait a moment."
[1354] 4. Sending response: The server converts the response into voice data and sends it back to the intercom to be conveyed to the delivery person.
[1355] 5. User notification: The server notifies the user of the visitor information and the response to their mobile device. The user receives a notification that a delivery person has visited them. Response: "Hello, it's you, a delivery person. Please wait a moment."
[1356] 6. Accepting user response: The user opens the app and selects how to respond. Since they are busy, they select automatic response.
[1357] 7. Execute Auto-Reply (Away Mode): If the user has Away Mode enabled, a message will be generated saying "I'm out of the office right now. Please come back later."
[1358] 8. Learning and risk assessment: The server learns from the video and audio of the delivery person and performs risk assessment. It determines that the person is not suspicious.
[1359] 9. Execute Shutout: Send a shutout message depending on conditions such as absence or risk of suspicious person. Not necessary this time.
[1360] 10. Analysis by emotion engine: The emotions (e.g., anxiety) that users feel when interacting with visitors are analyzed.
[1361] 11. Adjusting responses based on emotion: If the server is stressed, it will generate a gentler response such as "Please wait a moment."
[1362] For example, the prompt to input to a generative AI model might look like this:
[1363] Caller's voice: "This is a courier."
[1364] Prompt to generative AI: "The visitor says 'This is a courier.' Please generate an appropriate first-level response."
[1365] This allows the system to provide highly efficient visitor support.
[1366] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1367] Step 1: Signal reception
[1368] When a visitor presses the intercom, the server receives the signal via the Internet. This signal contains the visitor's video and audio data. The input is the signal when the visitor presses the intercom, and the output is the video and audio data stored on the server. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server, and the server captures the video and audio at that moment.
[1369] Step 2: Data analysis
[1370] The server analyzes the received video and audio data. First, it converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text. Next, it analyzes the video data and recognizes the visitor's face. This converts the input audio data into text format and also generates facial recognition data. For example, the audio saying "This is a delivery person" is converted into the text "Takhaibindesu" and the visitor's face is identified.
[1371] Step 3: Response Generation
[1372] The server generates an appropriate first-order response using a generative AI model (e.g., OpenAI's GPT-4) based on the text data and facial recognition data. The input is the text-converted voice data and facial recognition data, and the output is the response text. This response text is then converted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[1373] Step 4: Send response
[1374] The server sends the generated voice data to the intercom, and the visitor can receive the generated response. The input is the voice data generated by the server, and the output is the voice response transmitted to the visitor through the intercom. For example, the voice transmitted to the delivery person might say, "Hello, you are a delivery person. Please wait a moment."
[1375] Step 5: Notification
[1376] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The input is the visitor information (face image, part of the voice text) and the generated response, and the output is a notification message. A notification containing the visitor's face image and text is displayed on the user's mobile device. For example, the notification may read, "A delivery person has visited. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[1377] Step 6: Accept user response
[1378] The device displays the received notification for the user to review. The user can then review it and choose to respond manually, have it automatically respond, or shut it out. The input is the notification message from the server, and the output is the user's response decision. For example, if the user is busy and chooses to have it automatically respond, they can select it through the app.
[1379] Step 7: Auto-responders
[1380] If the user is in away mode, the server automatically generates an appropriate response and sends it to the intercom. The input is the away mode setting, and the output is an automatic response message. For example, a message like "I'm not available right now. Please come back later" is generated.
[1381] Step 8: Learning and risk assessment
[1382] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. The input is past visitor data and current visitor data, and the output is the suspicious person risk assessment result. If the server determines there is a high risk, it sends a warning to the user.
[1383] Step 9: Shut Out
[1384] The user can select the shut-out option on the mobile app. In this case, the server generates a shut-out message and transmits it to the visitor through the intercom. The input is the user's shut-out instruction, and the output is the shut-out message. For example, a message such as "Please refrain from visiting" is generated.
[1385] Step 10: Emotion Engine Analysis
[1386] The server analyzes the user's video and audio data and recognizes the user's emotions. The input is the user's video and audio data, and the output is the recognized emotional information. For example, it may be recognized from the user's video that they are feeling stressed.
[1387] Step 11: Adjust your response based on your emotions
[1388] The server adjusts the behavior of the notification and response generation means based on the recognized user emotion. The input is the recognized emotion information, and the output is a response adjusted according to the emotion. For example, if the user is feeling stressed, the generative AI model can generate a gentler tone of response or switch to automatic response mode.
[1389] (Application example 2)
[1390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1391] With conventional intercom systems, responding to visitors is done manually, which is not only time-consuming for users but also poses issues in terms of safety and efficiency. Even with electronic payment services, when a problem occurs during a transaction, the response is often delayed, causing frustration for users. Furthermore, insufficient evaluation of fraud risk leaves users with security concerns.
[1392] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means. This automates visitor response and transaction management, enabling safe and efficient responses while taking into consideration the user's emotional state.
[1393] The "signal receiving means" is a function that allows the server to receive signals from the interface system.
[1394] The "data analysis means" is a function for analyzing the video and audio data acquired from the interface system to identify user information.
[1395] The "response generation means" is a function for generating a primary response using a generative AI model based on the analyzed data.
[1396] The "response sending means" is a function for sending the generated response to the user through the interface system.
[1397] The "notification means" is a function for notifying the mobile terminal of the user information and the generated response content.
[1398] The "user response receiving means" is a function for receiving a response decision from a mobile terminal.
[1399] The "automatic response means" is a function for generating an automatic response based on the absence mode set by the user and sending it to the user.
[1400] The "learning and evaluation means" is a function for learning from past transaction data and evaluating fraud risk.
[1401] The "emotion analysis means" is a function for analyzing the emotional state of the user.
[1402] The "response adjustment means" is a function for adjusting the response based on the analysis results of the emotion analysis means.
[1403] The present invention provides a system for automating visitor response and transaction management, enabling safe and efficient response while taking into consideration the emotional state of the user. The system includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means.
[1404] System configuration and operation
[1405] 1. Signal receiving means:
[1406] The server receives a signal from an interface system (e.g., an electronic payment terminal) to obtain transaction data.
[1407] 2. Data analysis methods:
[1408] The server analyzes the video and audio data acquired from the interface system to identify the user's information. Specifically, it converts the audio data into text using a speech recognition system (e.g., Google Cloud Speech-to-Text) and analyzes it.
[1409] 3. Response Generation Method:
[1410] A generative AI model (e.g., GPT-4) is used to generate a first-order response based on the analyzed data. Example prompts for response generation are as follows:
[1411] "User is experiencing payment error. Stressed. Please generate an appropriate customer support response."
[1412] 4. Response sending method:
[1413] The generated response is converted into text or speech and sent to the user through the interface system, using text-to-speech (TTS) technology (e.g., Amazon Polly) to reconvert the speech.
[1414] 5. Means of notification:
[1415] The server then sends the user information and the generated response to the mobile device (smartphone) using a real-time notification system such as Firebase Cloud Messaging.
[1416] 6. User response acceptance means:
[1417] The mobile terminal provides an interface for receiving the notification and accepting the user's response. The user can respond manually or select an automatic response.
[1418] 7. Automated Response Methods:
[1419] Based on the out-of-office mode set by the user, an automated response is generated and sent to the user. The out-of-office response is constructed by the generative AI and includes a message informing the user that they are out of the office.
[1420] 8. Learning and assessment tools:
[1421] The server has an algorithm that learns from past transaction data and evaluates fraud risk. If the fraud risk is assessed as high, it will send a warning to the user.
[1422] 9. Emotion analysis means:
[1423] The server uses an emotion engine (e.g., IBM Watson or Microsoft Azure Emotion API) to analyze the user's emotional state, detecting emotions such as stress, anxiety, and joy.
[1424] 10. Response Adjustment Measures:
[1425] The server adjusts the response based on the analysis results of the emotion analysis means, and if the user is feeling stressed, it generates a response with a more reassuring tone.
[1426] Examples:
[1427] For example, consider a scenario in which a user encounters a payment error while using an electronic payment service and seeks support. A signal is sent from the interface system to the server, which converts the voice data into text. The server generates an appropriate response, converts it into voice, and sends it to the user. At the same time, the mobile device is notified of the details of the problem and the generated response. Furthermore, the server uses an emotion engine to analyze the user's emotional state and adjust the response accordingly. An example of a specific prompt would be, "The user is experiencing a payment error. He is in a stressful state. Please generate an appropriate customer support response."
[1428] This system reduces the burden on users in dealing with visitors and managing transactions, allowing them to receive safe and efficient services. In addition, the introduction of an emotion engine makes it possible to respond according to the user's emotional state, providing even greater peace of mind and convenience.
[1429] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1430] Step 1:
[1431] Signal Reception
[1432] The server receives signals from the interface system (e.g., electronic payment terminal). These signals include visitor information and transaction data. The input is the signal from the interface system, and the output is the received signal data. The signal sent from the interface system is received by the server's API.
[1433] Step 2:
[1434] Data analysis
[1435] The server analyzes the received video and audio data. The audio data is first converted into text using a speech recognition system (e.g., Google Cloud Speech-to-Text). Analysis is performed based on this text data and video data. The input is the received signal data, and the output is text data and analyzed video data. The audio data is converted into text format and then compared with the video data.
[1436] Step 3:
[1437] Response Generation
[1438] The server generates a primary response using a generative AI model (e.g., GPT-4) based on the analyzed data. The input is text data and video data, and the output is the text data of the generated primary response. For example, a response is generated based on the prompt sentence, "The user is experiencing a payment error. He is in a stressful state. Please generate an appropriate customer support response."
[1439] Step 4:
[1440] Response Send
[1441] The generated response is converted from text to speech. Text-to-speech (TTS) technology (e.g., Amazon Polly) is used for the speech conversion. The response is then sent to the user through the interface system via the response sending means. The input is the text data of the generated response, and the output is audio data. The text data is converted to audio data and transmitted to the user through the interface system.
[1442] Step 5:
[1443] notification
[1444] The server notifies the mobile device of the user's information along with the generated response. Notifications are made using a real-time notification system such as Firebase Cloud Messaging. The input is the response and user information, and the output is notification data sent to the mobile device. The notification system sends notifications to the user's smartphone in real time.
[1445] Step 6:
[1446] User response reception
[1447] The mobile terminal receives the notification and provides an interface for the user to respond or select an automatic response. The input is the notification data and the output is the user's selection. The user can review the response and select either a manual or automatic response.
[1448] Step 7:
[1449] Auto-response generation
[1450] The server generates an automatic response based on the away mode set by the user. The away response is constructed by a generation AI and includes a message informing the user that they are away. The input is the away mode status, and the output is the auto-response message. An away auto-response is generated and sent to the visitor.
[1451] Step 8:
[1452] Risk Assessment
[1453] The server studies past transaction data and evaluates fraud risk. If the risk is assessed as high, it issues a warning to the user. The input is past transaction data, and the output is the fraud risk assessment result and a warning notification. Past data is analyzed by an algorithm to evaluate risk.
[1454] Step 9:
[1455] Emotion analysis
[1456] The server uses an emotion engine to analyze the user's emotional state. The input is the user's video and audio data, and the output is the emotion analysis results. For example, it can detect the user's stress or anxiety.
[1457] Step 10:
[1458] Response Adjustment
[1459] The server adjusts the response based on the emotion analysis results. If the user is feeling stressed, it generates a response with a tone that gives a sense of security. The input is the emotion analysis result, and the output is the adjusted response. A response with an appropriate tone according to the emotion is generated and sent to the user.
[1460] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1461] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1462] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1463] [Fourth embodiment]
[1464] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1465] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1466] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1467] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1468] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1470] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1471] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1472] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1473] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1474] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1475] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1476] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1477] The present invention is a system for automating safe and efficient visitor attendance related to an intercom system. The system is implemented using the following components:
[1478] System configuration
[1479] 1. Signal receiving means:
[1480] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[1481] 2. Data analysis methods:
[1482] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[1483] 3. Response Generation Method:
[1484] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[1485] 4. Response sending method:
[1486] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[1487] 5. Means of notification:
[1488] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[1489] 6. User response acceptance means:
[1490] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether to respond to it themselves.
[1491] 7. Automated Response Methods:
[1492] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[1493] 8. Learning and assessment tools:
[1494] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[1495] 9. Shut-out measures:
[1496] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[1497] Explanation of program processing
[1498] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[1499] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[1500] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[1501] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[1502] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[1503] This system reduces the burden on users in dealing with visitors and allows them to communicate with them safely and efficiently.
[1504] The processing flow will be explained below.
[1505] Step 1:
[1506] A visitor presses the intercom
[1507] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[1508] Step 2:
[1509] Sending data to the server
[1510] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[1511] Step 3:
[1512] First-order response by generative AI
[1513] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[1514] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[1515] The server uses the generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[1516] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[1517] Step 4:
[1518] Notification of visitor information to mobile app
[1519] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[1520] Step 5:
[1521] User response decision
[1522] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[1523] The user can then decide whether to respond based on that information.
[1524] Step 6:
[1525] User response
[1526] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[1527] The user can connect to the intercom using the app and begin talking directly with the visitor.
[1528] Step 7:
[1529] Out of Office Replies
[1530] The server will automatically generate the appropriate response if the user has configured unattended mode.
[1531] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[1532] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[1533] Step 8:
[1534] Learn visitor characteristics and identify suspicious individuals
[1535] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[1536] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[1537] If the risk score is high, the server will notify the user with a warning.
[1538] Step 9:
[1539] Shutdown execution
[1540] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[1541] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[1542] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[1543] Example 1
[1544] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1545] In today's society, where safety and convenience are essential, the automation of visitor responses is becoming increasingly important. However, with conventional intercom systems, responses are performed manually, which places a burden on busy users. Furthermore, when a suspicious person visits, an immediate and appropriate response is required, but manual responses can be risky. Therefore, there is an urgent need to develop a system that automates visitor responses over intercoms in a safe and efficient manner.
[1546] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1547] In this invention, the server includes receiving means for receiving intercom signals, analyzing means for analyzing video and audio data acquired from the intercom and identifying visitor information, generating means for generating a primary response using a generative AI model based on the analyzed data, transmitting means for sending the generated response to the visitor via the intercom, notifying means for notifying a mobile device of the visitor information and the generated response content, receiving means for accepting a response decision from the mobile device, automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, evaluation means for learning visitor characteristics and comparing them with past data to evaluate suspicious person risk, and means for notifying the mobile device of the evaluated visitor information. This enables automated visitor response, safe and efficient visitor management, and rapid response to suspicious person risks.
[1548] The "receiving means" is a device or program for receiving an intercom signal and acquiring video and audio data.
[1549] The "analysis means" is a device or program that has the function of analyzing the video and audio data acquired by the receiving means and identifying information about the visitor.
[1550] The "generation means" is a device or program that generates a primary response using a generative AI model based on the data analyzed by the analysis means.
[1551] The "transmitting means" is a device or program that transmits the response generated by the generating means to the visitor via the intercom.
[1552] The "notification means" is a device or program for notifying the mobile terminal of the visitor information and the generated response content.
[1553] The "accepting means" is a device or program that has the function of accepting a response decision from a mobile terminal.
[1554] An "automatic response means" is a device or program that has the function of generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[1555] The "assessment means" is a device or program that has the function of learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[1556] A "generative AI model" is an artificial intelligence algorithm or platform that generates appropriate responses based on analyzed data.
[1557] A "prompt" is an instruction given to a generative AI model that serves as a basis for generating a specific response.
[1558] The present invention provides an improved intercom system for automating visitor responses. The system includes a receiving unit, an analyzing unit, a generating unit, a transmitting unit, a notifying unit, a reception unit, an automatic response unit, and an evaluation unit.
[1559] Hardware and Software
[1560] 1. Receiving means
[1561] Hardware: Intercom, Server
[1562] Software: The receiving program receives signals from the intercom via the Internet and acquires video and audio data.
[1563] 2. Analysis method
[1564] Hardware: Server
[1565] Software: Uses voice recognition systems (e.g., Google Speech-to-Text) to convert voice data into text, and runs image processing algorithms (e.g., facial recognition technology) to analyze video data.
[1566] 3. Generation means
[1567] Hardware: Server
[1568] Software: Generative AI models (e.g., OpenAI GPT-3) are used to generate a first-order response based on the analyzed text and video data.
[1569] 4. Transmission Method
[1570] Hardware: Servers, intercoms
[1571] Software: Software that converts the response text into speech (e.g., a text-to-speech engine) and a transmitting program that sends it to the intercom.
[1572] 5. Means of notification
[1573] Hardware: Servers, mobile devices
[1574] Software: A program that uses the push notification function of the mobile app to notify the mobile device of visitor information and the generated response.
[1575] 6. Method of reception
[1576] Hardware: Mobile devices
[1577] Software: A mobile application that provides a user interface where users can view the response and choose whether to respond manually or via an automated response.
[1578] 7. Automated Response Methods
[1579] Hardware: Server
[1580] Software: A program that generates an automatic response message based on the absence mode set by the user, converts it into voice data, and sends it to the intercom.
[1581] 8. Evaluation Methods
[1582] Hardware: Server
[1583] Software: Machine learning algorithms that learn visitor characteristics and programs that assess risk of suspicious behavior based on historical data.
[1584] Specific processing flow
[1585] First, when a visitor presses the intercom button, the signal is sent to the server. The receiving means receives this and captures the video and audio data. Next, the server's analysis means converts the audio data into text using a voice recognition system, and analyzes the video data using facial recognition technology.
[1586] The generating means uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate primary response based on the text data and video data. This generated response is output as text data, converted into audio data by the transmitting means, and transmitted to the visitor via the intercom. Through this process, the visitor can receive the response generated by the AI.
[1587] The notification means also transmits the analysis results to the mobile device so that the user can check the response. The user can choose to respond manually using the reception means or leave it to the automatic response means. If the absence mode is set, the automatic response means automatically generates and transmits an appropriate response.
[1588] Furthermore, the evaluation means can learn visitor data, evaluate the risk of suspicious individuals, and issue a warning to the user as necessary.
[1589] Examples of specific examples and prompts
[1590] For example, if the visitor is a delivery person, the voice data of the delivery person saying "Delivery service" is sent to the server. The server uses a voice recognition system to generate the text "Delivery service," and the generative AI model then generates a response such as "Hello, you are a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom.
[1591] Examples of prompts include:
[1592] "Please tell me how you would respond if the visitor was a delivery person."
[1593] Please explain how to respond if a suspicious person visits.
[1594] This reduces the burden on users in dealing with visitors and allows them to communicate with visitors safely and efficiently.
[1595] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1596] Step 1:
[1597] The visitor presses the intercom button.
[1598] Input: Visitor Action
[1599] Output: Signal from intercom to server
[1600] Specific operation: When the intercom button is pressed, video and audio data is generated and sent to the server along with a signal.
[1601] Step 2:
[1602] The server receives the signal transmitted from the interphone via the Internet.
[1603] Input: Intercom signal, video and audio data
[1604] Output: Video and audio data stored in a database
[1605] Specific operation: The server's receiving means captures the signal and stores the video and audio data in a database.
[1606] Step 3:
[1607] The server analyzes the received voice data and converts it into text using a voice recognition system (e.g., Google Speech-to-Text).
[1608] Input: Received audio data
[1609] Output: Text data
[1610] Specific operations: The server's analysis means analyzes the voice data using a voice recognition system and generates corresponding text data.
[1611] Step 4:
[1612] The server analyzes the received video data and attempts to identify the visitor using facial recognition technology.
[1613] Input: Received video data
[1614] Output: Visitor's specific information (face image, identification information)
[1615] What it does: The server's analytics runs the video data through image processing algorithms to recognize and identify the visitor's face and stores that information.
[1616] Step 5:
[1617] The server generates a first-order response based on the analyzed text and video data using a generative AI model (e.g., OpenAI GPT-3).
[1618] Input: Text data, video data
[1619] Output: Text data of the primary responses
[1620] Specific operation: The server's generation means inputs text data and video data into the generative AI model and generates a first response such as "Hello, how can I help you?"
[1621] Step 6:
[1622] The server converts the generated primary response into voice data and transmits it to the intercom.
[1623] Input: Text data of the primary response
[1624] Output: Audio data, sent to intercom
[1625] Specific operation: The server's transmission means converts the primary response into voice data using a text-to-speech engine, and transmits the voice data to the intercom.
[1626] Step 7:
[1627] The server notifies the mobile terminal of the analysis results and the generated response.
[1628] Input: Visitor information, content of initial response
[1629] Output: Notification message to mobile device
[1630] Specific operation: The server's notification means uses a push notification service to send the visitor's video, part of the audio text, and the response content to the mobile device.
[1631] Step 8:
[1632] The mobile terminal provides an interface for the user to check the notification content and select a manual or automatic response.
[1633] Input: Notification message from the server
[1634] Output: User response selection (manual or automatic)
[1635] What it does: Displays an interface on the mobile device that allows the user to view the notification and then provide options for how to respond.
[1636] Step 9:
[1637] If the user has enabled away mode, the server automatically generates an appropriate response and sends it to the visitor.
[1638] Input: User's away mode setting
[1639] Output: Auto-answer message, sent to intercom
[1640] Specific operation: The server's automatic response means uses the generative AI model to generate a message such as "I'm not here right now," converts it into voice data, and sends it to the intercom.
[1641] Step 10:
[1642] The server learns the visitor's characteristic data and compares it with past data to assess the risk of suspicious behavior.
[1643] Input: Visitor video and audio data, past visitor data
[1644] Output: Risk assessment results, warning notifications if necessary
[1645] Specific operation: The server's evaluation means uses a machine learning algorithm to analyze visitor data, assess the risk of suspicious activity, and if the risk is high, issue a warning to the user.
[1646] This detailed processing step ensures visitor interaction is automated, secure, and efficient.
[1647] (Application example 1)
[1648] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1649] With conventional intercom systems, communication with visitors is done manually, which places a heavy burden on the person in charge and makes it difficult to respond, especially when the person in charge is not present. Furthermore, in busy environments such as logistics centers, efficient visitor response is required. Furthermore, it is difficult to assess the risk of suspicious individuals, and measures to ensure safety are insufficient. To solve these issues, a system that automates visitor response safely and efficiently is needed.
[1650] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1651] In this invention, the server includes a signal receiving means for receiving an intercom signal, a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information, a response generation means for generating a primary response using a generative AI model based on the analyzed data, a response transmission means for sending the generated response to the visitor via the intercom, a notification means for notifying a mobile device of the visitor information and the generated response, a user response receiving means for receiving a response decision from the mobile device, an automatic response means for generating an automatic response based on an absence mode set by the user and sending it to the visitor, a learning and evaluation means for learning visitor characteristics and comparing them with past data to evaluate the risk of suspicious activity, a means for converting visitor information into text using a voice recognition system, a means for generating a primary response based on the text data using a generative AI model, a means for converting the generated text response into voice data, a communication means for notifying a smartphone or tablet of the analyzed visitor information, and a means for providing a user interface for determining a response based on the notified information. This reduces the burden on visitors and enables safe and efficient visitor response.
[1652] The "signal receiving means" is a means for receiving a signal from the interphone and processing the signal.
[1653] The "data analysis means" is a means for analyzing the video and audio data acquired from the intercom and identifying the visitor's information.
[1654] A "response generation means" is a means for generating a primary response using a generative AI model based on the analyzed data.
[1655] The "response transmitting means" is a means for transmitting the generated response to the visitor via the intercom.
[1656] The "notification means" is a means for notifying the mobile terminal of the visitor information and the generated response content.
[1657] The "user response receiving means" is a means for receiving a response decision from a mobile terminal.
[1658] The "automatic response means" is a means for generating an automatic response based on the absence mode set by the user and sending it to the visitor.
[1659] The "learning and evaluation means" is a means for learning the characteristics of visitors and evaluating the risk of suspicious persons by comparing them with past data.
[1660] A "voice recognition system" is a system for converting visitor information from voice data into text.
[1661] A "generative AI model" is an artificial intelligence model for generating first-order responses based on text data.
[1662] "Speech conversion means" refers to means for converting the generated text response into voice data.
[1663] "Communication means" refers to the means for notifying analyzed visitor information to a smartphone or tablet.
[1664] A "user interface" is a means for providing an interface for determining a response based on notified information.
[1665] To put the present invention into practice, the following system and its operation will be described. This system automates visitor reception at a logistics center in a safe and efficient manner.
[1666] First, a signal receiving means is provided to receive an intercom signal. This signal receiving means serves to acquire video and audio data generated by pressing the intercom button. Next, a data analysis means analyzes the video and audio data acquired from the intercom and identifies visitor information. The identified information is processed by a response generation means that generates a primary response using a generative AI model based on the analyzed data.
[1667] The generated primary response is sent to the visitor via the intercom via the response sending means. At the same time, the visitor information and the generated response are sent to the administrator's mobile device via the notification means. The mobile device is assumed to be a smartphone or tablet.
[1668] The mobile terminal is provided with a user response reception means, and the administrator can decide whether to respond manually or continue with an automatic response based on the received notification. The response decision may also be made automatically based on the absence mode set by the user. In this case, the automatic response means generates an appropriate response and sends it to the visitor.
[1669] In addition, the server has a learning and evaluation means that learns the characteristics of visitors and compares them with past data to evaluate the risk of suspicious persons. If the risk of suspicious persons is evaluated as high, a warning is sent to the user and, if necessary, a shut-out instruction is executed. A shut-out message is sent to the visitor via the intercom.
[1670] The system also includes a means for converting visitor information into text using a speech recognition system. For example, the Google Cloud Speech-to-Text API is used to convert the visitor's speech into text. A primary response based on the text data is generated using a generative AI model (e.g., OpenAI GPT-4). The generated text response is converted into voice data using a speech conversion means (e.g., Amazon Polly).
[1671] The server also has a communication method for notifying the smartphone or tablet of the analyzed visitor information. This communication is performed using a service such as Firebase Cloud Messaging (FCM). A UI framework such as React Native is used to provide a user interface for determining a response based on the notified information.
[1672] For example, if a visitor says, "Today's delivery," this voice data is converted into text data, "Today's delivery," via the Google Cloud Speech-to-Text API. A generative AI model receives this text data and generates a response, for example, "Hello, you're the delivery person. Please wait a moment." This response is converted into voice data by Amazon Polly and conveyed to the visitor via the intercom. At the same time, this information is also notified to the administrator's smartphone.
[1673] Example prompt: "Generate an appropriate response if the visitor says, 'Delivery today.'"
[1674] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1675] Step 1:
[1676] When the interphone button is pressed, the signal receiving means receives a signal from the interphone. This signal includes video data and audio data. After receiving the signal, these data are transmitted to the server.
[1677] Input: Intercom signal
[1678] Output: Video data, audio data
[1679] Step 2:
[1680] The server analyzes the received video and audio data using a data analysis tool. The audio data is converted to text using the Google Cloud Speech-to-Text API. The video data is analyzed using a facial recognition algorithm to obtain visitor information.
[1681] Input: Video data, audio data
[1682] Output: Text data, visitor information (face recognition results)
[1683] Step 3:
[1684] The server generates a first response using a generative AI model based on the analyzed text data and visitor information. OpenAI GPT-4 is used for this. The generated first response is output as text data.
[1685] Input: Text data, visitor information
[1686] Output: Text data of the primary response
[1687] Step 4:
[1688] The server converts the generated text data of the primary response into voice data using Amazon Polly as a voice conversion means.
[1689] Input: Text data of the primary response
[1690] Output: Audio data of the first response
[1691] Step 5:
[1692] The server transmits the generated voice data of the primary response to the intercom through the response transmitting means, thereby notifying the visitor of the response.
[1693] Input: Primary response audio data
[1694] Output: Visitor receives a voice response
[1695] Step 6:
[1696] The server notifies the administrator of the analyzed visitor information and the generated primary response content via a notification means to the administrator's smartphone or tablet.
[1697] Input: Visitor information, temporary response text data
[1698] Output: Notification to the administrator's smartphone
[1699] Step 7:
[1700] When the administrator receives a notification on their smartphone or tablet, the visitor's video, audio, and initial response are displayed, allowing the user to decide whether to respond manually or select an automatic response.
[1701] Input: Notified information (visitor's video, audio text, initial response)
[1702] Output: User response prompt
[1703] Step 8:
[1704] If the user judges that the visitor is a high risk of being a suspicious person, the server uses learning and evaluation means to evaluate the risk, and if necessary, issues a warning to the administrator and executes a shut-out instruction. A shut-out message is sent to the visitor via the intercom.
[1705] Input: Information about suspicious person risk, user shut-out instructions
[1706] Output: Shutout message
[1707] Step 9:
[1708] The server analyzes and learns from visitor characteristics and past data via learning and evaluation means, and reflects this in future responses, thereby improving the accuracy of future visitor risk assessments.
[1709] Input: Visitor characteristics, historical data
[1710] Output: Learning results, updated evaluation model
[1711] In this way, visitor reception can be automated, safely, and efficiently achieved.
[1712] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1713] The present invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components.
[1714] System configuration
[1715] 1. Signal receiving means:
[1716] It receives intercom signals and captures video and audio data. When a visitor presses the intercom, the signal is sent to the server in the system.
[1717] 2. Data analysis methods:
[1718] The server receives the video and audio data sent from the intercom. This data is first converted into text by a voice recognition system, and then compared with the video data.
[1719] 3. Response Generation Method:
[1720] The server uses a generative AI model based on the text and video data to generate an appropriate initial response to the visitor. The response is output as text, which is then converted back into voice data and transmitted to the visitor via the intercom.
[1721] 4. Response sending method:
[1722] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[1723] 5. Means of notification:
[1724] The server then sends the analyzed visitor information and the generated response to the user's mobile device. The notification includes the visitor's facial image, a portion of the voice text, and the response.
[1725] 6. User response acceptance means:
[1726] The mobile device displays the received notification in a form that the user can view, and an interface is provided for the user to view the notification and decide whether or not to respond to it.
[1727] 7. Automated Response Methods:
[1728] When a user activates away mode, the server automatically generates an appropriate response and sends it to the visitor through the intercom. The away response is constructed by a generative AI and includes messages such as "I'm not available right now."
[1729] 8. Learning and assessment tools:
[1730] The server analyzes the visitor's video and audio data and learns their characteristics. This is used to compare the visitor's past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, a warning is sent to the user.
[1731] 9. Shut-out measures:
[1732] Users can select the shut-out option in the mobile app, in which case the server generates a shut-out message and transmits it to the visitor over the intercom.
[1733] 10. Emotion Engine:
[1734] The server is equipped with an emotion engine that analyzes the user's video and audio data to recognize the user's emotions. The emotion engine analyzes the emotions (e.g., stress, anxiety, joy, surprise, etc.) that the user shows while interacting with visitors and adjusts the system's response based on that information.
[1735] 11. Emotion-based response adjustment measures:
[1736] The server adjusts the operation of the notification means and response generation means based on the user's emotions recognized by the emotion engine. For example, if the user is feeling stressed or anxious, the response generation means generates a response with a gentler tone or switches to an automatic response.
[1737] Explanation of program processing
[1738] The server has a dedicated algorithm that analyzes the video and audio data sent from the intercom and generates an automatic response. First, when a visitor presses the intercom, the signal is sent to the server via the Internet. This signal contains the visitor's video and audio data.
[1739] The server analyzes the received data and converts the audio into text. Next, a generative AI model generates an appropriate initial response based on the text and video data. The generated response is output as text data, which is then converted back into audio data and sent over the intercom. This process allows the visitor to receive a response from the AI.
[1740] The server also sends the analyzed visitor information to the mobile device. The user can check the notification received on the mobile device and choose to respond manually, have it automatically respond, or shut out the call.
[1741] For example, if the visitor is a delivery person, they can say "Delivery service here" into the intercom, and the voice data will be sent to the server. The server analyzes the voice to obtain the text data "Delivery service," and the generative AI model generates a primary response: "Hello, you're a delivery person. Please wait a moment." This response is converted into voice data and transmitted to the delivery person via the intercom. At the same time, the user's mobile device will be notified of the delivery person's video and text, allowing them to decide whether to respond themselves.
[1742] Furthermore, the server analyzes the user's video and audio using an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed, the emotion engine detects this and the server adjusts the response generation means based on that information. As a result, the generated response is delivered in an appropriate tone according to the user's emotional state.
[1743] On the other hand, if the visitor is likely to be suspicious, the server uses past data to assess the risk and sends a warning to the user. After receiving the warning, the user can choose to shut the visitor out by selecting the shut-out option on their mobile device. In this case, a shut-out message generated by the server is transmitted to the visitor via the intercom.
[1744] This system reduces the burden on users when dealing with visitors and allows them to communicate with them safely and efficiently. In addition, the introduction of an emotion engine enables responses that take into account the user's emotional state, providing even greater peace of mind and convenience.
[1745] The processing flow will be explained below.
[1746] Step 1:
[1747] A visitor presses the intercom
[1748] When a visitor presses the intercom, the intercom receives the signal and begins recording video and audio.
[1749] Step 2:
[1750] Sending data to the server
[1751] The intercom transmits recorded video and audio data to a server in real time via the Internet.
[1752] Step 3:
[1753] First-order response by generative AI
[1754] The server analyzes the received video and audio data and converts the visitor's voice into text (voice recognition).
[1755] The server uses the text and video data to instruct the generative AI model to generate a primary response.
[1756] The server uses a generative AI model to generate an appropriate response and converts that response into voice data (text-to-speech synthesis).
[1757] The server transmits the generated voice data to an intercom via the Internet and responds to the visitor.
[1758] Step 4:
[1759] Notification of visitor information to mobile app
[1760] The server notifies the mobile app of the visitor's video, audio, and generated response data. The notification data includes the visitor's facial image, a portion of the voice text, and the response data.
[1761] Step 5:
[1762] User response decision
[1763] The user receives a push notification on their mobile app and opens the app to view the visitor's information, including video, audio, transcribed voice recordings, and generated responses.
[1764] The user can then decide whether to respond based on that information.
[1765] Step 6:
[1766] User response
[1767] If the user decides to respond, he or she presses the "Reply" button on the mobile app.
[1768] The user can connect to the intercom using the app and begin talking directly with the visitor.
[1769] Step 7:
[1770] Out of Office Replies
[1771] The server will automatically generate the appropriate response if the user has configured unattended mode.
[1772] The server uses a generative AI model to generate an automatic response such as "I'm not here right now" and converts it into audio data.
[1773] The server transmits the generated voice data to the intercom and conveys it to the visitor.
[1774] Step 8:
[1775] Learn visitor characteristics and identify suspicious individuals
[1776] The server stores the recorded video and audio data and uses machine learning models to learn the characteristics of visitors, analyzing their facial recognition and voice patterns.
[1777] The server compares the data with past data to assess the risk of suspicious activity, and if suspicious activity is detected, calculates a risk score.
[1778] If the risk score is high, the server will notify the user with a warning.
[1779] Step 9:
[1780] Shutdown execution
[1781] When a user receives a warning notification, they can select the shut-out option in the mobile app.
[1782] Based on the user's selection, the server generates a shut-out message for the visitor and converts it into voice data.
[1783] The server transmits the generated shut-out voice data to the intercom and conveys it to the visitor.
[1784] Step 10:
[1785] Recognizing user emotions with an emotion engine
[1786] The server acquires the user's video and audio data from the user's mobile terminal.
[1787] The server uses an emotion engine to analyze the user's video and audio data and recognize the user's emotions.
[1788] Step 11:
[1789] Emotion-Based Response Modulation
[1790] The server adjusts the response generation means based on the user's emotions recognized by the emotion engine.
[1791] If the user shows signs of stress or anxiety, the server generates a response in a gentle tone using the response generating means, or switches to an automatic response.
[1792] Example 2
[1793] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1794] Conventional intercom systems require users to respond to visitors manually, placing a heavy burden on users. It is also difficult to respond appropriately when the resident is away or to dangerous visitors, resulting in safety and convenience issues. Furthermore, the system is unable to respond flexibly to the user's emotional state, which can be stressful.
[1795] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a signal receiving means for receiving an intercom signal; a data analysis means for analyzing video and audio data acquired from the intercom and identifying visitor information; a response generation means for generating a primary response using a generative model based on the analyzed data; a response transmission means for transmitting the generated response to the visitor via the intercom; a notification means for notifying the mobile communication device of the visitor information and the generated response content; a user response receiving means for receiving a response decision from the mobile communication device; an automatic response means for generating an automatic response based on an absence mode set by the user and transmitting it to the visitor; a learning and evaluation means for learning visitor characteristics and evaluating the risk of a suspicious person by comparing them with past data; a means for analyzing the user's video and audio data and having an emotion engine that recognizes the user's emotions; and an emotion-based response adjustment means for adjusting the operation of the notification means and the response generation means based on the recognized user emotions. This automates visitor response, reducing the burden on the user and enabling appropriate responses to be taken when the user is absent or when a dangerous visitor is present. Furthermore, a system that can respond flexibly according to the user's emotional state is provided, allowing the system to be used with peace of mind.
[1796] The "signal receiving means" is a device or function for receiving a signal transmitted from the intercom.
[1797] "Data analysis means" refers to a device or function that analyzes the video and audio data acquired from the intercom and identifies visitor information.
[1798] A "response generator" is a device or function that uses a generative model based on analyzed data to generate a primary response.
[1799] The "response transmitting means" is a device or function for transmitting the generated response to the visitor via the intercom.
[1800] The "notification means" refers to a device or function for notifying the portable communication device of the visitor's information and the generated response content.
[1801] The "user response receiving means" is a device or function for receiving a response decision from the portable communication device.
[1802] The "automatic response means" is a device or function that generates an automatic response based on the absence mode set by the user and sends it to the visitor.
[1803] "Learning and evaluation means" refers to devices and functions that learn the characteristics of visitors and compare them with past data to evaluate the risk of suspicious persons.
[1804] An "emotion engine" is a device or function that analyzes a user's video and audio data and recognizes the user's emotions.
[1805] The "emotion-based response adjustment means" is a device or function for adjusting the operation of the notification means or response generation means based on the recognized user emotion.
[1806] This invention is a system for automating visitor responses related to intercoms in a safe and efficient manner. In particular, by combining it with an emotion engine that recognizes the user's emotions, more advanced visitor responses can be achieved. This system is implemented using the following components:
[1807] System configuration
[1808] 1. Signal receiving means:
[1809] The server receives the intercom signal and captures the video and audio data. When a visitor presses the intercom, the signal is sent to the server via the Internet. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server.
[1810] 2. Data analysis methods:
[1811] The server receives the video and audio data sent from the intercom. It converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text, and then compares it with the video data. For example, a voice saying "This is a delivery" can be converted into text like "Takhaibindesu."
[1812] 3. Response Generation Method:
[1813] The server generates an appropriate response using a generative AI model (e.g., OpenAI's GPT-4) based on the text and video data from the voice. The generated response is first output in text format and then reconverted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[1814] 4. Response sending method:
[1815] The server transmits the generated voice data to an intercom, where the visitor can receive the generated response.
[1816] 5. Means of notification:
[1817] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The notification includes an image of the visitor's face, a portion of the voice text, and the response. For example, the user's smartphone may receive a notification saying, "A delivery person has visited you. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[1818] 6. User response acceptance means:
[1819] The mobile device displays the received notification for the user to review. The user can then review the notification and choose to respond to it themselves, have it automatically respond, or shut it out. For example, if the user is busy, they can choose to automatically respond.
[1820] 7. Automated Response Methods:
[1821] If the user is in away mode, the server automatically generates an appropriate response to send to the intercom, such as "I'm not here right now. Please come back later."
[1822] 8. Learning and assessment tools:
[1823] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. If the risk is high, it sends a warning to the user.
[1824] 9. Shut-out measures:
[1825] Users can select the shut-out option on their mobile app, in which case the server will generate a shut-out message and send it to the visitor over the intercom, such as "Please refrain from visiting."
[1826] 10. Emotion Engine:
[1827] The server is equipped with an emotion engine that analyzes the user's video and audio data and recognizes the user's emotions, such as stress, anxiety, joy, surprise, etc.
[1828] 11. Emotion-based response adjustment measures:
[1829] The server adjusts the content and tone of the generative AI model's responses based on the user's perceived emotions. For example, if the user is feeling stressed, the server can generate responses in a gentler tone or switch to an automatic response mode.
[1830] Examples of specific examples and prompts
[1831] For example, consider a scenario where a delivery person arrives.
[1832] 1. Signal reception: The delivery person presses the intercom, and a "ping pong" signal is sent to the server.
[1833] 2. Analysis of audio and video data: The server converts the audio "This is a delivery" into the text "Takhaibindesu" and recognizes the visitor's face.
[1834] 3. Generate a response: The generative AI model (e.g., GPT-4) generates a response such as, "Hello, you're a delivery person. Please wait a moment."
[1835] 4. Sending response: The server converts the response into voice data and sends it back to the intercom to be conveyed to the delivery person.
[1836] 5. User notification: The server notifies the user of the visitor information and the response to their mobile device. The user receives a notification that a delivery person has visited them. Response: "Hello, it's you, a delivery person. Please wait a moment."
[1837] 6. Accepting user response: The user opens the app and selects how to respond. Since they are busy, they select automatic response.
[1838] 7. Execute Auto-Reply (Away Mode): If the user has Away Mode enabled, a message will be generated saying "I'm out of the office right now. Please come back later."
[1839] 8. Learning and risk assessment: The server learns from the video and audio of the delivery person and performs risk assessment. It determines that the person is not suspicious.
[1840] 9. Execute Shutout: Send a shutout message depending on conditions such as absence or risk of suspicious person. Not necessary this time.
[1841] 10. Analysis by emotion engine: The emotions (e.g., anxiety) that users feel when interacting with visitors are analyzed.
[1842] 11. Adjusting responses based on emotion: If the server is stressed, it will generate a gentler response such as "Please wait a moment."
[1843] For example, the prompt to input to a generative AI model might look like this:
[1844] Caller's voice: "This is a courier."
[1845] Prompt to generative AI: "The visitor says 'This is a courier.' Please generate an appropriate first-level response."
[1846] This allows the system to provide highly efficient visitor support.
[1847] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1848] Step 1: Signal reception
[1849] When a visitor presses the intercom, the server receives the signal via the Internet. This signal contains the visitor's video and audio data. The input is the signal when the visitor presses the intercom, and the output is the video and audio data stored on the server. For example, when a visitor presses the "ping-dong" button, the signal is sent to the server, and the server captures the video and audio at that moment.
[1850] Step 2: Data analysis
[1851] The server analyzes the received video and audio data. First, it converts the audio data into text using a speech recognition system such as Google Cloud Speech-to-Text. Next, it analyzes the video data and recognizes the visitor's face. This converts the input audio data into text format and also generates facial recognition data. For example, the audio saying "This is a delivery person" is converted into the text "Takhaibindesu" and the visitor's face is identified.
[1852] Step 3: Response Generation
[1853] The server generates an appropriate first-order response using a generative AI model (e.g., OpenAI's GPT-4) based on the text data and facial recognition data. The input is the text-converted voice data and facial recognition data, and the output is the response text. This response text is then converted into voice data using Google Text-to-Speech. For example, a response might be generated that says, "Hello, you're a delivery person. Please wait a moment."
[1854] Step 4: Send response
[1855] The server sends the generated voice data to the intercom, and the visitor can receive the generated response. The input is the voice data generated by the server, and the output is the voice response transmitted to the visitor through the intercom. For example, the voice transmitted to the delivery person might say, "Hello, you are a delivery person. Please wait a moment."
[1856] Step 5: Notification
[1857] The server notifies the user's mobile device of the analyzed visitor information and the generated response. The input is the visitor information (face image, part of the voice text) and the generated response, and the output is a notification message. A notification containing the visitor's face image and text is displayed on the user's mobile device. For example, the notification may read, "A delivery person has visited. Response: 'Hello, you are a delivery person. Please wait a moment.'"
[1858] Step 6: Accept user response
[1859] The device displays the received notification for the user to review. The user can then review it and choose to respond manually, have it automatically respond, or shut it out. The input is the notification message from the server, and the output is the user's response decision. For example, if the user is busy and chooses to have it automatically respond, they can select it through the app.
[1860] Step 7: Auto-responders
[1861] If the user is in away mode, the server automatically generates an appropriate response and sends it to the intercom. The input is the away mode setting, and the output is an automatic response message. For example, a message like "I'm not available right now. Please come back later" is generated.
[1862] Step 8: Learning and risk assessment
[1863] The server analyzes the visitor's video and audio data and learns their characteristics. This allows it to compare them with past visitor patterns and evaluate the risk of a suspicious person. The input is past visitor data and current visitor data, and the output is the suspicious person risk assessment result. If the server determines there is a high risk, it sends a warning to the user.
[1864] Step 9: Shut Out
[1865] The user can select the shut-out option on the mobile app. In this case, the server generates a shut-out message and transmits it to the visitor through the intercom. The input is the user's shut-out instruction, and the output is the shut-out message. For example, a message such as "Please refrain from visiting" is generated.
[1866] Step 10: Emotion Engine Analysis
[1867] The server analyzes the user's video and audio data and recognizes the user's emotions. The input is the user's video and audio data, and the output is the recognized emotional information. For example, it may be recognized from the user's video that they are feeling stressed.
[1868] Step 11: Adjust your response based on your emotions
[1869] The server adjusts the behavior of the notification and response generation means based on the recognized user emotion. The input is the recognized emotion information, and the output is a response adjusted according to the emotion. For example, if the user is feeling stressed, the generative AI model can generate a gentler tone of response or switch to automatic response mode.
[1870] (Application example 2)
[1871] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1872] With conventional intercom systems, responding to visitors is done manually, which is not only time-consuming for users but also poses issues in terms of safety and efficiency. Even with electronic payment services, when a problem occurs during a transaction, the response is often delayed, causing frustration for users. Furthermore, insufficient evaluation of fraud risk leaves users with security concerns.
[1873] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means. This automates visitor response and transaction management, enabling safe and efficient responses while taking into consideration the user's emotional state.
[1874] The "signal receiving means" is a function that allows the server to receive signals from the interface system.
[1875] The "data analysis means" is a function for analyzing the video and audio data acquired from the interface system to identify user information.
[1876] The "response generation means" is a function for generating a primary response using a generative AI model based on the analyzed data.
[1877] The "response sending means" is a function for sending the generated response to the user through the interface system.
[1878] The "notification means" is a function for notifying the mobile terminal of the user information and the generated response content.
[1879] The "user response receiving means" is a function for receiving a response decision from a mobile terminal.
[1880] The "automatic response means" is a function for generating an automatic response based on the absence mode set by the user and sending it to the user.
[1881] The "learning and evaluation means" is a function for learning from past transaction data and evaluating fraud risk.
[1882] The "emotion analysis means" is a function for analyzing the emotional state of the user.
[1883] The "response adjustment means" is a function for adjusting the response based on the analysis results of the emotion analysis means.
[1884] The present invention provides a system for automating visitor response and transaction management, enabling safe and efficient response while taking into consideration the emotional state of the user. The system includes a signal receiving means, a data analyzing means, a response generating means, a response sending means, a notification means, a user response accepting means, an automatic response means, a learning and evaluation means, an emotion analysis means, and a response adjustment means.
[1885] System configuration and operation
[1886] 1. Signal receiving means:
[1887] The server receives a signal from an interface system (e.g., an electronic payment terminal) to obtain transaction data.
[1888] 2. Data analysis methods:
[1889] The server analyzes the video and audio data acquired from the interface system to identify the user's information. Specifically, it converts the audio data into text using a speech recognition system (e.g., Google Cloud Speech-to-Text) and analyzes it.
[1890] 3. Response Generation Method:
[1891] A generative AI model (e.g., GPT-4) is used to generate a first-order response based on the analyzed data. Example prompts for response generation are as follows:
[1892] "User is experiencing payment error. Stressed. Please generate an appropriate customer support response."
[1893] 4. Response sending method:
[1894] The generated response is converted into text or speech and sent to the user through the interface system, using text-to-speech (TTS) technology (e.g., Amazon Polly) to reconvert the speech.
[1895] 5. Means of notification:
[1896] The server then sends the user information and the generated response to the mobile device (smartphone) using a real-time notification system such as Firebase Cloud Messaging.
[1897] 6. User response acceptance means:
[1898] The mobile terminal provides an interface for receiving the notification and accepting the user's response. The user can respond manually or select an automatic response.
[1899] 7. Automated Response Methods:
[1900] Based on the out-of-office mode set by the user, an automa...
Claims
1. a signal receiving means for receiving an intercom signal; a data analysis means for analyzing the video and audio data acquired from the intercom and identifying information about the visitor; a response generation means for generating a primary response using a generative AI model based on the analyzed data; a response transmitting means for transmitting the generated response to the visitor through the intercom; a notification means for notifying the mobile terminal of the visitor's information and the generated response content; a user response accepting means for accepting a response decision from the mobile terminal; an automatic response means for generating an automatic response based on an absence mode set by the user and sending the automatic response to the visitor; A learning and evaluation method that learns the characteristics of visitors and compares them with past data to evaluate the risk of suspicious individuals; A system including:
2. means for issuing a warning to a user when the learning and evaluation means evaluates that the user is at high risk of being a suspicious person, and for issuing a shut-out instruction as necessary; 10. The system of claim 1, further comprising means for transmitting a shut-out message to the visitor through an intercom.
3. The system according to claim 1, further comprising means for displaying the visitor's video, audio and response content when the mobile terminal receives the notification, and providing an interface for the user to respond or select an automatic response.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A