system

The system automates intercom interactions through sensor detection, generative AI dialogue, speech recognition, and natural language processing to handle visitor requests, enhancing user convenience and safety.

JP2026062281APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current intercom systems require direct user interaction, which can be inconvenient when the user is absent or busy, posing safety and convenience risks, especially in situations like remote meetings or dealing with unwanted visitors.

Method used

A system utilizing sensors to detect visitors, generative artificial intelligence for initial dialogue, speech recognition to convert speech to text, natural language processing to classify requests, and automated response generation, with notifications sent to the user's device.

Benefits of technology

Automates visitor interactions, reducing user burden and improving safety and convenience by handling responses efficiently and appropriately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062281000001_ABST
    Figure 2026062281000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A sensor means for detecting visitors, A server means that receives notifications from the aforementioned sensor means, activates a generative artificial intelligence, and performs an initial dialogue. The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text, A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose, A response means for generating an automated response message according to the classified request and communicating it to the visitor, A notification means for notifying the user's terminal of the aforementioned request and response content, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the current intercom system, since the user needs to respond directly, there is a problem that it is difficult to respond when absent or busy. Also, when only children are left at home or during a remote meeting, etc., in situations where an immediate response is not possible, there is a high risk that safety and convenience will be impaired. Furthermore, dealing with unwanted visitors such as persistent solicitations and sales is also troublesome for the user. Due to these problems, the user's life may become complicated and safety may be impaired.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides a system that includes a sensor means for detecting visitors, a server means for receiving notifications from the sensor means and activating a generative artificial intelligence to conduct an initial dialogue, a speech recognition means for the generative artificial intelligence to convert speech into text, a natural language processing means for analyzing the text and classifying the visitor's request, a response means for generating an automatic response message according to the classified request and communicating it to the visitor, and a notification means for notifying the user's terminal of the request and response content. This system can automatically respond when the user is away, during remote meetings, when children are home alone, and even to persistent solicitations and salespeople, thereby improving user safety and convenience.

[0006] A "sensor" refers to a device that detects the movements of visitors and transmits that information to a server.

[0007] A "server system" refers to a computer system that receives notifications from sensor systems, activates generative artificial intelligence to process that information, and has the function of conducting initial dialogue.

[0008] "Generative artificial intelligence" refers to AI programs that possess the ability to formulate hypotheses and engage in dialogue, just like humans.

[0009] "Voice recognition means" refers to an algorithm that converts a visitor's voice into text data.

[0010] "Natural language processing means" refers to technology that analyzes text data generated by speech recognition means and appropriately classifies the visitor's purpose.

[0011] "Response means" refers to a device or system that has the function of generating an appropriate automated response message based on the requirements classified by natural language processing means and communicating it to the visitor.

[0012] "Notification means" refers to devices or systems that have communication functions to inform the user's terminal of the request and response content.

[0013] "User's device" refers to portable devices such as smartphones and tablets that users use on a daily basis. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] The system for carrying out this invention automatically processes a series of operations from visitor detection to response and notification. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, and notification means.

[0036] Initial processing when the intercom rings

[0037] Terminal:

[0038] When a visitor arrives at the intercom, a sensor detects them and sends an event signal to the server. This allows the server to recognize that a visitor has arrived.

[0039] server:

[0040] Upon receiving the event signal, the server activates the generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[0041] Dialogue with visitors

[0042] Terminal:

[0043] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[0044] server:

[0045] The server converts the received audio data into text data using speech recognition. The converted text data is then analyzed by natural language processing to classify the visitor's request.

[0046] Classification of inquiries and automated responses

[0047] server:

[0048] Natural language processing tools analyze text data and classify visitors' requests into several categories. These categories include, for example, unattended delivery, solicitation or sales calls, remote meetings, and children home alone. An automated response message tailored to each request is then generated by a generative artificial intelligence system.

[0049] Terminal:

[0050] The generated response message is transmitted to the visitor through the device's speaker, ensuring they receive an appropriate response.

[0051] Notification to the user

[0052] server:

[0053] The generated response and information regarding the visitor's purpose are sent as push notifications to the user's device using a notification system. For example, if the user has a smartphone, a specific notification such as "We have confirmed that Sagawa Express has arrived and left the package at your front door" will be sent.

[0054] User:

[0055] Users can review received notifications, generate a follow-up message if necessary, and transmit it to visitors via the server. This follow-up allows users to send custom messages, such as "I'll be right back, please wait a moment," when a friend visits.

[0056] Specific example

[0057] Here are some specific examples.

[0058] Example 1: Leaving delivery unattended when the recipient is absent.

[0059] 1. Terminal: The intercom rings.

[0060] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0061] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver a package."

[0062] 4. Server: Receives voice data, converts it to text, and analyzes it using natural language processing. Determines that it is a delivery attempt when the recipient is absent.

[0063] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0064] 6. Terminal: Provides instructions to visitors.

[0065] 7. Server: Notifies the user that "A delivery service has arrived. The package has been left in front of your door."

[0066] Example 2: Dealing with persistent solicitation / sales tactics

[0067] 1. Terminal: The intercom rings.

[0068] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0069] 3. Terminal: The visitor responds, "I've come to introduce a new service."

[0070] 4. Server: Receives audio data and converts it to text. Analysis using natural language processing techniques indicates that it is a solicitation / sales pitch.

[0071] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0072] 6. Terminal: Provides instructions to visitors.

[0073] This system can streamline intercom communication in users' daily lives, improving security and convenience.

[0074] The following describes the processing flow.

[0075] Step 1:

[0076] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[0077] Step 2:

[0078] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[0079] Step 3:

[0080] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[0081] Step 4:

[0082] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[0083] Step 5:

[0084] Terminal: Sends recorded audio data to the server via the internet.

[0085] Step 6:

[0086] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[0087] Step 7:

[0088] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[0089] Step 8:

[0090] Server: Based on the classification results, an appropriate automated response message is generated by the AI ​​generation system. This message includes the optimal response method for the given request.

[0091] Step 9:

[0092] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[0093] Step 10:

[0094] Terminal: The response message is properly conveyed to the visitor, and the response ends.

[0095] Step 11:

[0096] Server: Notifies the user's terminal of the generated response and information regarding the visitor's purpose.

[0097] Step 12:

[0098] Users can receive notifications on devices such as smartphones, and, if necessary, generate additional instructions or response messages via the server and communicate them to visitors.

[0099] These steps automate the intercom system, allowing for efficient handling of visitors without direct user interaction.

[0100] (Example 1)

[0101] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] Traditional intercom systems require users to manually handle a series of processes, from detecting visitors to answering, providing appropriate responses, and notifying the user, which places a significant burden on the user. This inconvenience is particularly noticeable when users are away for extended periods or working remotely. Furthermore, dealing with inappropriate solicitations and sales calls presents challenges, highlighting the need for automated systems to address these issues.

[0103] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0104] In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to conduct an initial dialogue, a speech recognition means in which the generative artificial intelligence converts speech into text, a natural language processing means for analyzing the text and classifying the visitor's request, a response means for generating an automated response message according to the classified request and communicating it to the visitor, and a notification means for notifying the user's communication device of the request and response content. As a result, a series of processes from visitor detection to response and notification are automated, significantly reducing the burden on the user, improving convenience, and ensuring security.

[0105] A "sensor device" is a device that detects visitors, generates an event signal, and notifies the server.

[0106] "Generative artificial intelligence" is a program that generates text questions and answers to automate interactions with visitors.

[0107] "Speech recognition means" refers to a device or program that converts received speech data into text data.

[0108] A "natural language processing tool" is a program that analyzes text data, understands its meaning, and classifies it into appropriate categories.

[0109] A "response means" refers to a device or program for transmitting a generated response message to an visitor.

[0110] A "notification means" is a device or program that notifies the user's communication device of the generated response content or information regarding the visitor's purpose.

[0111] A "server system" is a computer system that activates generative artificial intelligence and controls and manages various systems in a centralized manner.

[0112] "Communication devices" refer to terminals that allow users to receive notifications, and include smartphones and tablets.

[0113] The system for implementing this invention automatically processes a series of actions from visitor detection to response and notification. The hardware includes sensors, terminals, servers, and communication equipment, while the software includes generative artificial intelligence, a speech recognition engine, a natural language processing engine, and a notification system.

[0114] First, the system uses sensors to detect visitors. For example, infrared sensors or motion sensors may be used. When a visitor is detected, the sensor generates an event signal, which is transmitted to the server via the network.

[0115] When the server receives this event signal, it activates a generative artificial intelligence (e.g., the GPT-3® model). The generative AI then generates prompt sentences to start the initial conversation. For example, it might generate questions such as "Who is this?", "What can I help you with?", or "Who in your family is this addressed to?". The generated prompt sentences are then converted into speech data. A text-to-speech (TTS) engine (e.g., Google® Text-to-Speech API) performs this conversion.

[0116] The generated audio data is played back through the terminal's speaker, conveying the question to the visitor. Next, the intercom's microphone records the visitor's response in real time and sends the audio data to the server. The server uses speech recognition (e.g., IBM Watson® Speech to Text API) to convert the received audio data into text data.

[0117] The converted text data is analyzed using natural language processing tools (e.g., the spaCy library), and the visitor's request is classified into several categories. For example, possible categories include receiving packages when absent, refusing solicitations, responding during remote meetings, and handling children left alone at home.

[0118] Based on the classification results, the generative artificial intelligence generates an appropriate response message. For example, a message such as, "I am currently out, please leave your package in front of the door." This response message is sent to the terminal and conveyed to the visitor through the intercom speaker. The generated response content and the visitor's request are recorded on the server and also pushed to the user's communication device using a notification method (e.g., Firebase Cloud Messaging).

[0119] Users can check notifications on communication devices such as smartphones and generate a follow-up message if necessary. By sending this follow-up message to the server, it becomes possible to provide visitors with appropriate instructions again.

[0120] Specific example

[0121] Example 1: Leaving delivery unattended when the recipient is absent.

[0122] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[0123] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[0124] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver your package," and sends the voice data to the server.

[0125] 4. Server: Uses speech recognition to convert speech data into text data, which is then analyzed using natural language processing. This is classified as "delivery left unattended when the recipient is absent."

[0126] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0127] 6. Terminal: Receives a response message from the visitor via the speaker.

[0128] 7. Server: Notifies the user's communication device with the message, "A delivery service has arrived. The package has been left in front of your door."

[0129] Example 2: Dealing with persistent solicitation / sales tactics

[0130] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[0131] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[0132] 3. Terminal: The visitor responds, "I've come to introduce you to a new service," and sends the audio data to the server.

[0133] 4. Server: Converts speech data into text data using speech recognition technology and analyzes it using natural language processing technology. Classified as solicitation / sales.

[0134] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0135] 6. Terminal: Receives a response message from the visitor via the speaker.

[0136] This system allows users to significantly automate their intercom responses, improving the convenience and safety of their daily lives.

[0137] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0138] Step 1:

[0139] The terminal detects visitors using the intercom's sensor. The sensor consists of infrared sensors and motion sensors, and generates an event signal when it detects a visitor's movement. The input is the presence of a visitor, and the output is the transmission of the event signal to the server.

[0140] Step 2:

[0141] The server receives an event signal from a sensor. Upon receiving the event signal, the server activates a generative artificial intelligence (e.g., the GPT-3 model) to generate a prompt for initial interaction. The input is the event signal, and the output is the prompt. Specifically, the generative artificial intelligence generates a question such as "Who are you?"

[0142] Step 3:

[0143] The server uses a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) to convert the generated prompt text into speech data. The input is the prompt text, and the output is speech data. The TTS engine outputs the prompt text as synthesized speech.

[0144] Step 4:

[0145] The terminal receives audio data transmitted from the server and plays it back through the intercom speaker. The input is audio data, and the output is the audio message conveyed to the visitor. The speaker plays the audio data and conveys the question to the visitor.

[0146] Step 5:

[0147] The terminal uses the intercom's microphone to record the visitor's response in real time. The recorded audio data is immediately sent to the server. The input is the visitor's voice, and the output is the transmission of audio data to the server.

[0148] Step 6:

[0149] The server converts the received audio data into text data using speech recognition software (e.g., IBM Watson Speech to Text API). The input is audio data, and the output is text data. The speech recognition software analyzes the audio and outputs it as text data.

[0150] Step 7:

[0151] The server sends text data to a natural language processing system (e.g., the spaCy library), which analyzes the content and classifies the visitor's request. The input is text data, and the output is the classification result. The natural language processing system performs analysis and classification, assigning the request to a specific category (e.g., receiving a package when absent, refusing a solicitation, etc.).

[0152] Step 8:

[0153] The server uses generative artificial intelligence to generate an appropriate response message based on the classification result. The input is the classification result, and the output is the response message. The generative artificial intelligence then generates a prompt sentence, which is used as the response message.

[0154] Step 9:

[0155] The terminal receives a response message (audio data) from the server and transmits it to the visitor through the intercom speaker. The input is the response message, and the output is the audio transmitted to the visitor. The speaker plays the response message and responds to the visitor.

[0156] Step 10:

[0157] The server records the response and visitor's request and sends a push notification to the user's device using a notification method (e.g., Firebase Cloud Messaging). The input is the response and visitor's request, and the output is a push notification to the user's device. The notification method generates the message and sends it to the user's device.

[0158] Step 11:

[0159] The user can review the notification and, if necessary, generate a reply message and send it to the server. The input is the user's reply message, and the output is the message sent to the server. The user inputs and sends the message using a smartphone or similar device.

[0160] Step 12:

[0161] The server processes the received reply message and transmits it to the visitor through the terminal. The input is the reply message, and the output is the audio message to be delivered to the visitor. The server converts the reply message into audio data and plays it through the terminal's speaker.

[0162] The above outlines the specific processing steps of this system. This automates the entire process from visitor detection to response and notification, significantly reducing the burden on the user.

[0163] (Application Example 1)

[0164] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0165] In recent years, many homes and businesses have frequently had to deal with visitors, and prompt and appropriate responses are especially required when people are away or busy. However, current intercom systems require manual responses, which is inefficient and poses security risks. Furthermore, it is difficult to respond quickly to unexpected visitors such as solicitors or salespeople, resulting in wasted time and effort for users. Moreover, if these responses are not handled properly, important notifications may be delayed, which can negatively impact users' quality of life.

[0166] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0167] In this invention, the server includes detection means for detecting visitors, computation means for activating generative artificial intelligence to conduct initial dialogue, speech recognition means for converting speech to text, natural language processing means for analyzing text and classifying the visitor's request, response means for generating and responding with an automated response message, speech synthesis means for generating a voice response, and push notification means for sending visitor information to the user's terminal. This makes it possible to automate visitor interactions efficiently and securely. Furthermore, it enables appropriate responses in various situations such as handling deliveries when the user is absent and refusing solicitations, reducing the burden on the user and improving security.

[0168] A "visitor" is a person who visits a specific location, such as a home or office, and attempts to make contact using an intercom.

[0169] A "detection means" is a device or sensor used to detect the presence of a visitor, and it plays the first role in recognizing a visitor.

[0170] "Generative artificial intelligence" refers to artificial intelligence technologies that automatically perform intelligent tasks such as conversation, text generation, and speech recognition.

[0171] "Computation means" refers to a part of a server or computer system that receives and processes data from sensors and other devices.

[0172] "Voice recognition means" refers to a technology or device that converts voice input into text data, and is a means of recording the content of conversations with visitors as text.

[0173] "Natural language processing" refers to technologies that classify and understand specific intentions and requirements through the analysis of text data.

[0174] A "response means" is a technology or device for generating an appropriate response message based on classified requirements and communicating it to visitors.

[0175] A "speech synthesis engine" is a technology that converts text data into speech data and outputs it as speech.

[0176] "Push notification methods" refer to technologies for sending notifications to a user's device in real time.

[0177] This invention is a system that automatically processes a series of actions from visitor detection to response and notification. This system mainly consists of the following components: a detection means for detecting visitors, a computation means for activating generative artificial intelligence and conducting an initial dialogue, a speech recognition means for converting speech to text, a natural language processing means for analyzing text data and classifying the visitor's request, a response means for generating a response message and communicating it to the visitor, a push notification means for sending visitor information to the user's terminal, and a speech synthesis engine.

[0178] When the server receives a signal from the sensor, it activates generative artificial intelligence and begins an initial conversation with the visitor. For example, it might ask questions such as, "Who is this? What can I help you with? Who in your family is this for?" The initial conversation is generated by a speech synthesis engine and output through the speaker. The visitor's responses are captured by the intercom microphone and converted into text through speech recognition.

[0179] Next, the server analyzes the converted text using natural language processing to classify the visitor's request. Based on the classified request, a generative artificial intelligence generates an appropriate response message and transmits it to the visitor via a response system.

[0180] For example, if a visitor responds, "It's a delivery service. I've come to deliver your package," the server generates a response message saying, "You are currently out, please leave your package in front of the door," and outputs it using a speech synthesis engine.

[0181] Furthermore, the server uses push notifications to inform the user's device of visitor information and response content. This notification is delivered in real time using APIs such as LINE Notify API. Users receive notifications through their devices and can generate custom response messages as needed, then transmit them to the visitor again via the server.

[0182] The following scenarios are possible as specific examples.

[0183] When a visitor presses the intercom, the following notification is sent to the smartphone: "Who is it? What can I do for you? Who in the family is this for?". If the visitor replies, "It's a delivery. I've come to deliver a package," the system responds, "You are currently out, please leave the package in front of the door," and simultaneously sends a notification to the user's smartphone saying, "A delivery person has arrived. They have left the package in front of the door."

[0184] In this way, this invention automates visitor reception, significantly reducing the burden on users and improving security.

[0185] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0186] Step 1:

[0187] The sensor detects the visitor.

[0188] Specific operation: The detection device (sensor) detects when a visitor presses the intercom and sends an event signal to the server. The input data is information about the visitor's presence, and the output data is the event signal sent to the server.

[0189] Step 2:

[0190] The server receives an event signal and activates the generative artificial intelligence.

[0191] Specific operation: The server receives an event signal from the sensor and activates a generative artificial intelligence (AI). The generative AI prepares initial questions and uses a speech synthesis engine to output a voice message such as, "Who is this? What can I help you with? Who in your family is this for?" The input data is the event signal, and the output data is a voice message to the visitor.

[0192] Step 3:

[0193] The device captures the visitor's response and sends it to the server.

[0194] Specific operation: The intercom's microphone captures the visitor's response and sends it to the server as audio data. The input data is the visitor's voice, and the output data is the audio data sent to the server.

[0195] Step 4:

[0196] The server receives the audio data and converts it into text using speech recognition technology.

[0197] Specific operation: The server converts audio data into text data using speech recognition technology. The input data is audio data, and the output data is the converted text data.

[0198] Step 5:

[0199] The server analyzes text data using natural language processing to classify the visitor's purpose.

[0200] Specific operation: The server analyzes text data using natural language processing and classifies it into categories such as "delivery" and "sales." Input data is text data, and output data is classified category information.

[0201] Step 6:

[0202] The server uses generative artificial intelligence to generate response messages based on classified requests, and a speech synthesis engine creates voice messages.

[0203] Specific operation: The server uses generative artificial intelligence to generate appropriate response messages for classified requests and outputs them as voice messages using a speech synthesis engine. For example, if a visitor says "I have a delivery," the server will generate a message saying, "I am currently out, please leave your package in front of the door." The input data is classified request information, and the output data is the response voice message.

[0204] Step 7:

[0205] The terminal transmits the generated response message to the visitor.

[0206] Specific operation: The generated response message is played to the visitor through the device's speaker. The input data is the response voice message, and the output data is the voice output to the visitor.

[0207] Step 8:

[0208] The server creates and sends a push notification to the user's device containing visitor information and response details.

[0209] Specific operation: The server uses push notification methods, such as the LINE Notify API, to send visitor information and response content as push notifications to the user's smartphone. The input data consists of the response content and visitor information, and the output data is the push notification message sent to the user's device.

[0210] Step 9:

[0211] The user generates a follow-up message as needed and sends it to the server.

[0212] Specific operation: The user receives a push notification on their smartphone, enters a reply message as needed, and sends it to the server. The input data is the user's reply message, and the output data is the reply data sent to the server.

[0213] Step 10:

[0214] The server receives a response message from the user, converts it into an appropriate format using generative artificial intelligence, and transmits it to the visitor via the terminal.

[0215] Specific operation: The server receives a response message from the user, converts it into a voice message using generative artificial intelligence, and plays it back to the visitor through the terminal's speaker. The input data is the user's response message, and the output data is the voice message to the visitor.

[0216] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0217] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[0218] Initial processing when the intercom rings

[0219] Terminal:

[0220] When a visitor arrives at the intercom, a sensor detects this and sends an event signal to the server. This allows the server to recognize the visitor's presence.

[0221] server:

[0222] Upon receiving the event signal, the server activates the generative artificial intelligence and initiates an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[0223] Dialogue with visitors

[0224] Terminal:

[0225] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[0226] server:

[0227] The received audio data is converted into text data using speech recognition technology. This text data accurately reflects the content of the visitor's speech.

[0228] server:

[0229] The converted text data is analyzed using natural language processing to classify the visitor's purpose into multiple categories. Examples of these categories include: unattended delivery, solicitation, remote meeting, and childcare. Based on the analysis results, an appropriate automated response message is generated by a generative artificial intelligence system.

[0230] Automated responses and user notifications

[0231] server:

[0232] The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. This ensures that the visitor receives an appropriate response.

[0233] server:

[0234] The generated response and information regarding the visitor's purpose are sent to the user's device using a notification system.

[0235] User emotion recognition and adjustment

[0236] server:

[0237] An emotion engine built into the user's device recognizes emotions from the user's text or voice input. Based on this emotional state, notification content can be adjusted. For example, if the user is stressed, notifications will be delivered using softer language.

[0238] server:

[0239] Furthermore, the emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[0240] User response

[0241] User:

[0242] Users who receive notifications on devices such as smartphones can generate additional instructions or response messages via the server as needed. For example, if a friend comes to visit, they can send a custom message such as, "I'll be right back, please wait a moment."

[0243] Specific example

[0244] Example 1: Leaving delivery unattended when the recipient is absent.

[0245] 1. Terminal: The intercom rings.

[0246] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0247] 3. Terminal: The visitor responds, "Delivery. I've come to deliver your package."

[0248] 4. Server: Receives voice data and converts it to text. Analyzes it using natural language processing to determine whether to leave the package unattended when the recipient is absent.

[0249] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0250] 6. Terminal: Provides instructions to visitors.

[0251] 7. Server: Notifies the user that "Your delivery has arrived. The package has been left in front of your door."

[0252] 8. Server: If the emotion engine detects the user's stress level, it adjusts the notification content to a softer tone.

[0253] Example 2: Solicitation / Sales Strategies

[0254] 1. Terminal: The intercom rings.

[0255] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0256] 3. Terminal: The visitor responds, "This is to inform you about a new service."

[0257] 4. Server: Receives audio data, converts it to text, analyzes it using natural language processing, and determines whether it is a solicitation / sales call.

[0258] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0259] 6. Terminal: Provides instructions to visitors.

[0260] 7. Server: Notifies the user that "a solicitation was received and declined."

[0261] 8. Server: The emotion engine dynamically adjusts the tone and content of notifications according to the user's situation.

[0262] In this way, it is possible to efficiently handle interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[0263] The following describes the processing flow.

[0264] Step 1:

[0265] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[0266] Step 2:

[0267] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[0268] Step 3:

[0269] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[0270] Step 4:

[0271] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[0272] Step 5:

[0273] Terminal: Sends recorded audio data to the server via the internet.

[0274] Step 6:

[0275] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[0276] Step 7:

[0277] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[0278] Step 8:

[0279] Server: Based on the analysis results, the AI ​​generates an appropriate automated response message. This message includes the most suitable response for the request.

[0280] Step 9:

[0281] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[0282] Step 10:

[0283] Terminal: The response message is properly conveyed to the visitor.

[0284] Step 11:

[0285] Server: Transmit the generated response content and information regarding the visitor's inquiry to the user's terminal using the notification means.

[0286] Step 12:

[0287] User: Receive the notification on a terminal such as a smartphone, and the emotion engine analyzes the user's emotion.

[0288] Step 13:

[0289] Server: The emotion engine recognizes the user's emotional state and adjusts the notification content as necessary. For example, when the user is feeling stressed, a notification with gentle language is sent.

[0290] Step 14:

[0291] Server: Based on the user's emotional state, the emotion engine can dynamically change the tone and content of the auto-response message. For example, when the user is relaxed, a message with a friendly tone is generated.

[0292] Step 15:

[0293] User: Check the notification, generate a re-response message from the smartphone if necessary, and convey it to the visitor through the server. For example, a custom message such as "Please wait a moment as I'll be back soon" can be sent.

[0294] Through these steps, this system can efficiently automate the response to the visitor while considering the user's emotional state, improving the user's safety and convenience.

[0295] (Example 2)

[0296] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0297] Traditional intercom systems require manual responses to visitors, making it difficult to provide appropriate care when the user is absent or busy. Furthermore, they lack the ability to respond flexibly to emotions and situations, hindering user convenience and safety. Additionally, they often fail to accurately identify the visitor's purpose and situation, leading to delays in responses.

[0298] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0299] In this invention, the server includes a sensor means for detecting visitors, a computing means for activating generative artificial intelligence to conduct initial dialogue, and a speech recognition means for converting speech to text. This enables automatic detection of visitors and appropriate initial dialogue. It also includes a natural language processing means for analyzing text and classifying the visitor's request, and an emotion engine for recognizing the user's emotions and adjusting notification content and responses. This enables flexible responses that take into account the user's emotional state, improving user safety and convenience.

[0300] A "visitor" is a person who visits via the intercom system.

[0301] "Sensor means" refers to devices or means used to detect the presence of visitors, and includes motion detection and infrared sensors.

[0302] "Computer means" refers to a system or device that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[0303] "Generative artificial intelligence" refers to artificial intelligence technology that generates voice dialogues and response messages within a system.

[0304] The "voice recognition means" is a technology or device for converting the voice of a visitor into text data.

[0305] The "natural language processing means" is a technology or device for analyzing text data and classifying the matters of a visitor.

[0306] The "response means" is a device or technology for generating an automatic response message according to the classified matter and transmitting it to the visitor.

[0307] The "notification means" is a technology or device for notifying the response content and the matters of a visitor to the user's terminal.

[0308] The "emotion engine" is a technology or system for recognizing the emotion of a user and adjusting the notification content and response.

[0309] The system for implementing this invention automatically performs a series of operations from detecting a visitor to responding and notifying, and further recognizes the emotion of a user to adjust the notification and response. This system includes sensor means, computer means, generative artificial intelligence, voice recognition means, natural language processing means, response means, notification means, and emotion engine.

[0310] Initial processing when the intercom rings

[0311] When the terminal detects a visitor, the sensor means detects this and transmits an event signal to the server. When the server receives the event signal, it activates the generative artificial intelligence and starts an initial conversation. The content of this conversation includes questions such as "Who is it?", "What is the matter?", and "Who is it addressed to in the family?". The generative artificial intelligence outputs this as voice using voice synthesis technology.

[0312] Conversation with the visitor

[0313] The device's microphone captures the visitor's responses and records them in real time. The recorded audio data is sent to a server. The server converts the received audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). After conversion, the text data is analyzed by a natural language processing unit (e.g., Hugging Face's NLP model) to classify the visitor's request into multiple categories.

[0314] Automated responses and user notifications

[0315] The server generates an appropriate automated response message based on the analysis results. The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. In addition, information regarding the response content and the visitor's purpose is sent to the user's terminal using notification methods (e.g., push notification to a smartphone or email).

[0316] User emotion recognition and adjustment

[0317] The server has an emotion engine that can recognize emotions from text and voice input from the user's device. The emotion engine utilizes the Emotion API and other tools to adjust notification content to a softer tone if the user is feeling stressed. The emotion engine can also dynamically adjust the tone and content of automated response messages based on the user's emotional state. If the user is relaxed, messages will be generated in a friendly tone.

[0318] User response

[0319] When a user receives a notification on a device such as a smartphone, they can generate additional instructions or reply messages via the server as needed. For example, if a friend visits, they can send a custom message such as, "I'll be right back, please wait a moment."

[0320] Specific example

[0321] Example 1: Leaving delivery unattended when the recipient is absent.

[0322] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[0323] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[0324] 3. Terminal: The intercom's microphone captures the visitor's speech, such as "Delivery. I've come to deliver your package," and sends it to the server.

[0325] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[0326] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if the package was left unattended when the recipient was absent.

[0327] 6. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0328] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[0329] 8. Server: Notifies the user, "The delivery has arrived. The package has been left in front of your door."

[0330] 9. Server: The server uses the Emotion API to sense the user's stress level and adjusts the notification content to a softer tone.

[0331] Example of a prompt:

[0332] "If your doorbell rings and a visitor asks you to leave a package, how would you respond?"

[0333] Example 2: Solicitation / Sales Strategies

[0334] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[0335] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[0336] 3. Terminal: The intercom's microphone captures the visitor's statement, "We have an announcement about a new service," and sends it to the server.

[0337] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[0338] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if it is a solicitation / sales pitch.

[0339] 6. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0340] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[0341] 8. Server: Notifies the user that "a solicitation was received and declined."

[0342] 9. Server: Use the Emotion API to check the user's status and dynamically adjust the tone and content of notifications.

[0343] In this way, the system efficiently handles interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[0344] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0345] Step 1:

[0346] Terminal: The intercom's sensor detects a visitor. As input, the sensor obtains information about the visitor's presence. When the sensor is triggered, an event signal is generated and sent to the server. As output, the server recognizes the visitor's presence.

[0347] Step 2:

[0348] Server: The server receives an event signal and activates the generative artificial intelligence. It receives the event signal as input. The generative artificial intelligence generates initial dialogue questions. It generates questions such as "Who is this?", "What can I help you with?", and "Who in your family is this addressed to?" as text and converts them into speech using speech synthesis technology. It generates the voiced questions as output and sends them to the terminal.

[0349] Step 3:

[0350] Terminal: The intercom's microphone captures and records the visitor's response. It receives the visitor's voice response as input. The recorded data is sent to the server. It also sends the voice data to the server as output.

[0351] Step 4:

[0352] Server: The server converts the received audio data into text data using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). It receives audio data as input. The speech recognition engine analyzes the audio and generates text data. It generates text data as output.

[0353] Step 5:

[0354] Server: The converted text data is analyzed using natural language processing (e.g., Hugging Face's NLP model). It receives text data as input. The natural language processing analyzes the text content and classifies the visitor's request into multiple categories. It generates classified category information as output.

[0355] Step 6:

[0356] Server: Generates an appropriate automated response message based on the analysis results. It receives classified category information as input. The generative artificial intelligence generates a response message such as, "We are currently out, please leave your package at the front door." It sends the generated response message to the terminal as output.

[0357] Step 7:

[0358] Terminal: The terminal transmits the response message received from the server to the visitor through the intercom speaker. Input: Receives the response message. Plays the voice message using the intercom speaker. Output: Transmits the response message to the visitor.

[0359] Step 8:

[0360] Server: The generated response and information regarding the visitor's request are sent to the user's device using a notification method. The server receives the response and request information as input. It then notifies the user's device using a notification method (e.g., push notification or email). The notification is then delivered to the user as output.

[0361] Step 9:

[0362] Server: The server contains an emotion engine that recognizes the user's emotions from text and voice input from the user's device. It receives user input data as input. It analyzes the user's emotional state using the Emotion API and other tools. It generates user emotion information as output.

[0363] Step 10:

[0364] Server: The emotion engine adjusts notification content based on the user's emotional state. It receives user emotion information as input. It adjusts the notification content and tone according to the emotional state to generate an appropriate notification message. As output, it sends the adjusted notification message to the user's device.

[0365] Step 11:

[0366] User: The user receives a notification on a device such as a smartphone and generates additional instructions or reply messages via the server as needed. Input: Receives the user notification. Creates a custom message (e.g., "I'll be right back, please wait a moment") as needed and sends it to the server. Output: The reply message is transmitted to the visitor.

[0367] (Application Example 2)

[0368] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0369] Traditional intercom systems often required manual interaction with visitors, making quick and appropriate responses difficult, especially when users were away from home, working remotely, or with children home alone. Furthermore, the lack of appropriate notifications and responses tailored to the user's emotional state often led to inconvenience and stress. As a result, many users felt a lack of both safety and convenience.

[0370] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to perform an initial dialogue, a means for the generative artificial intelligence to convert speech into text, a means for analyzing the text and classifying the visitor's request, a means for generating an automatic response message according to the classified request and transmitting it to the visitor, a means for notifying the request and response content to the user's communication terminal, and a means for recognizing the user's emotions and adjusting the tone and content of the notification and automatic response message. This enables quick and appropriate responses even when the user is absent or busy, and allows for flexible notifications and responses according to the user's emotional state.

[0371] "Sensor means" refers to a device or system used to detect visitors.

[0372] A "server means" is a device or system that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[0373] "Generative artificial intelligence" refers to artificial intelligence software that converts speech into text, analyzes that text, and generates appropriate response messages.

[0374] "Voice recognition means" refers to a device or system that converts a visitor's voice into text.

[0375] A "natural language processing device" is a device or system that analyzes text to classify the visitor's request.

[0376] "Response means" refers to a device or system that generates an automated response message according to the classified requirements and transmits it to the visitor.

[0377] "Notification means" refers to a device or system that notifies the user's communication terminal of the aforementioned request and response content.

[0378] "Emotion recognition means" refers to a device or system that recognizes the user's emotions and adjusts the tone and content of the notification and automated response messages.

[0379] A "communication terminal" is an electronic device, including smartphones and tablets, that a user uses to receive notifications.

[0380] The system for implementing this invention automatically performs a series of operations from visitor detection to response and notification, and further recognizes the user's emotions to adjust notifications and responses accordingly. The specific form of this system is described below.

[0381] Hardware configuration

[0382] This system uses the following hardware:

[0383] Sensory devices (e.g., intercoms, security cameras, door sensors)

[0384] Server configuration (e.g., cloud server)

[0385] User's communication device (e.g., smartphone, tablet)

[0386] Software to use

[0387] This system is equipped with the following software:

[0388] Generative artificial intelligence (e.g. GPT-3)

[0389] Speech recognition methods (e.g., Google Speech Recognition API)

[0390] Natural language processing tools (e.g., TextBlob library)

[0391] Emotion recognition methods (e.g., TextBlob's sentiment analysis function)

[0392] Data processing and data calculation

[0393] The server performs the following data processing and calculations.

[0394] 1. Visitor detection

[0395] Sensors detect the presence of visitors. Examples include intercoms and door sensors.

[0396] 2. Initiating the initial dialogue

[0397] The server receives a notification from a sensor and activates a generative artificial intelligence (AI) to initiate an initial conversation with the visitor, such as "Who are you?". This AI then uses speech recognition to convert the visitor's voice into text.

[0398] 3. Classification of Visitor's Purpose

[0399] Text data converted by generative artificial intelligence is analyzed using natural language processing techniques to appropriately classify the visitor's purpose.

[0400] 4. Generation and transmission of response messages

[0401] Based on the categorized request, an automated response message is generated and communicated to the visitor via the response system.

[0402] 5. Notification to the user's terminal

[0403] Simultaneously, the server sends information about the response and the visitor's purpose to the user's communication terminal using a notification mechanism.

[0404] 6. User emotion recognition and notification adjustment

[0405] Furthermore, using emotion recognition mechanisms, the system recognizes the user's emotions from their text or voice input and dynamically adjusts the tone of notification content and response messages.

[0406] Specific example

[0407] The following are specific examples.

[0408] Example 1: Delivery company's response

[0409] 1. The sensor detects the delivery person.

[0410] 2. The server activates a generative artificial intelligence and responds, "Who is this?"

[0411] 3. The delivery person responds, "I've come to deliver your package."

[0412] 4. The speech recognition means converts the speech into text, and the natural language processing means analyzes the text.

[0413] 5. The server generates a response message saying, "We are currently out, please leave it at the front door," and transmits it to the delivery person via the response system.

[0414] 6. The user's communication device is notified with the message, "Your delivery has arrived. The package has been left in front of your door."

[0415] Example 2: Detection of a suspicious person

[0416] 1. The sensor detects a suspicious person.

[0417] 2. The server activates a generative artificial intelligence and responds, "Who is it?", but the suspicious person remains silent.

[0418] 3. The server identifies the person as suspicious, generates a warning message, and transmits it via the response system.

[0419] 4. Simultaneously, a warning notification is sent to the user's communication terminal.

[0420] Example of a prompt

[0421] "Person says: I'm here to deliver a package. User is stressed. Generate a soft response."

[0422] In this way, it becomes possible to respond quickly and appropriately even when the user is away or busy, and to provide flexible notifications and responses that are tailored to their emotional state.

[0423] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0424] Step 1:

[0425] A sensor detects the presence of a visitor. The input is the sensor (e.g., intercom, security camera, door sensor), and the output is a notification signal of the detection event. This notification signal is sent to the server.

[0426] Step 2:

[0427] The server receives a notification from a sensor. The input is a notification signal, and the server activates a generative artificial intelligence to initiate an initial conversation. Specifically, it generates questions such as "Who are you?" for the visitor and outputs them as voice through a response system.

[0428] Step 3:

[0429] The visitor's response is captured by the intercom's microphone and recorded in real time as audio data. The input is the visitor's voice, and the output is the recorded audio data. The terminal sends this audio data to the server.

[0430] Step 4:

[0431] The server converts the received audio data into text data using speech recognition tools (e.g., Google Speech Recognition API). The input is audio data, and the output is text data. Specifically, the speech recognition engine analyzes the audio waveform data and converts it into the corresponding text.

[0432] Step 5:

[0433] The server analyzes the converted text data using natural language processing tools (e.g., the TextBlob library). The input is text data, and the output is the category of requests resulting from the analysis. In this process, the content of the text is analyzed, and its intent and requirements are classified into categories such as receiving packages when absent, refusing solicitations, responding during remote meetings, and how to handle situations when children are home alone.

[0434] Step 6:

[0435] The server generates an automated response message based on the analysis results. The input is the category of the request, and the output is the automated response message. Specifically, a generative artificial intelligence generates an appropriate message and conveys it to the visitor. For example, if the recipient is absent and the package is to be left at the door, the server will generate a message such as, "We are currently absent, please leave your package at the front door."

[0436] Step 7:

[0437] The server transmits a generated response message to the visitor via the response mechanism. The input is an automated response message, and the output is an audio message to the visitor. Specifically, it is transmitted to the visitor through the intercom speaker.

[0438] Step 8:

[0439] The server sends the generated response and information about the visitor's purpose to the user's communication terminal using a notification mechanism. The input is the response and purpose information, and the output is a notification message to the user's terminal. Specifically, it is displayed as a notification on the user's smartphone.

[0440] Step 9:

[0441] The server uses emotion recognition to recognize the user's emotions and adjusts the tone and content of notifications and automated response messages accordingly. The input is the user's emotion data (e.g., text or voice input), and the output is the adjusted notification and response message. For example, if the user is feeling stressed, the notification content might be adjusted to a softer tone, such as "Your package has been left at your front door."

[0442] Step 10:

[0443] The user generates additional instructions or follow-up messages as needed, which are then transmitted to the visitor via the server. The input is the user's follow-up message, and the output is a voice message to the visitor. For example, the user might send a custom message such as, "I'll be right back, please wait a moment."

[0444] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0445] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0446] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0447] [Second Embodiment]

[0448] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0449] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0450] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0451] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0452] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0454] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0455] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0456] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0457] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0458] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0459] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0460] The system for carrying out this invention automatically processes a series of operations from visitor detection to response and notification. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, and notification means.

[0461] Initial processing when the intercom rings

[0462] Terminal:

[0463] When a visitor arrives at the intercom, a sensor detects them and sends an event signal to the server. This allows the server to recognize that a visitor has arrived.

[0464] server:

[0465] Upon receiving the event signal, the server activates the generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[0466] Dialogue with visitors

[0467] Terminal:

[0468] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[0469] server:

[0470] The server converts the received audio data into text data using speech recognition. The converted text data is then analyzed by natural language processing to classify the visitor's request.

[0471] Classification of inquiries and automated responses

[0472] server:

[0473] Natural language processing tools analyze text data and classify visitors' requests into several categories. These categories include, for example, unattended delivery, solicitation or sales calls, remote meetings, and children home alone. An automated response message tailored to each request is then generated by a generative artificial intelligence system.

[0474] Terminal:

[0475] The generated response message is transmitted to the visitor through the device's speaker, ensuring they receive an appropriate response.

[0476] Notification to the user

[0477] server:

[0478] The generated response and information regarding the visitor's purpose are sent as push notifications to the user's device using a notification system. For example, if the user has a smartphone, a specific notification such as "We have confirmed that Sagawa Express has arrived and left the package at your front door" will be sent.

[0479] User:

[0480] Users can review received notifications, generate a follow-up message if necessary, and transmit it to visitors via the server. This follow-up allows users to send custom messages, such as "I'll be right back, please wait a moment," when a friend visits.

[0481] Specific example

[0482] Here are some specific examples.

[0483] Example 1: Leaving delivery unattended when the recipient is absent.

[0484] 1. Terminal: The intercom rings.

[0485] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0486] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver a package."

[0487] 4. Server: Receives voice data, converts it to text, and analyzes it using natural language processing. Determines that it is a delivery attempt when the recipient is absent.

[0488] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0489] 6. Terminal: Provides instructions to visitors.

[0490] 7. Server: Notifies the user that "A delivery service has arrived. The package has been left in front of your door."

[0491] Example 2: Dealing with persistent solicitation / sales tactics

[0492] 1. Terminal: The intercom rings.

[0493] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0494] 3. Terminal: The visitor responds, "I've come to introduce a new service."

[0495] 4. Server: Receives audio data and converts it to text. Analysis using natural language processing techniques indicates that it is a solicitation / sales pitch.

[0496] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0497] 6. Terminal: Provides instructions to visitors.

[0498] This system can streamline intercom communication in users' daily lives, improving security and convenience.

[0499] The following describes the processing flow.

[0500] Step 1:

[0501] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[0502] Step 2:

[0503] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[0504] Step 3:

[0505] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[0506] Step 4:

[0507] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[0508] Step 5:

[0509] Terminal: Sends recorded audio data to the server via the internet.

[0510] Step 6:

[0511] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[0512] Step 7:

[0513] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[0514] Step 8:

[0515] Server: Based on the classification results, an appropriate automated response message is generated by the AI ​​generation system. This message includes the optimal response method for the given request.

[0516] Step 9:

[0517] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[0518] Step 10:

[0519] Terminal: The response message is properly conveyed to the visitor, and the response ends.

[0520] Step 11:

[0521] Server: Notifies the user's terminal of the generated response and information regarding the visitor's purpose.

[0522] Step 12:

[0523] Users can receive notifications on devices such as smartphones, and, if necessary, generate additional instructions or response messages via the server and communicate them to visitors.

[0524] These steps automate the intercom system, allowing for efficient handling of visitors without direct user interaction.

[0525] (Example 1)

[0526] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0527] Traditional intercom systems require users to manually handle a series of processes, from detecting visitors to answering, providing appropriate responses, and notifying the user, which places a significant burden on the user. This inconvenience is particularly noticeable when users are away for extended periods or working remotely. Furthermore, dealing with inappropriate solicitations and sales calls presents challenges, highlighting the need for automated systems to address these issues.

[0528] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0529] In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to conduct an initial dialogue, a speech recognition means in which the generative artificial intelligence converts speech into text, a natural language processing means for analyzing the text and classifying the visitor's request, a response means for generating an automated response message according to the classified request and communicating it to the visitor, and a notification means for notifying the user's communication device of the request and response content. As a result, a series of processes from visitor detection to response and notification are automated, significantly reducing the burden on the user, improving convenience, and ensuring security.

[0530] A "sensor device" is a device that detects visitors, generates an event signal, and notifies the server.

[0531] "Generative artificial intelligence" is a program that generates text questions and answers to automate interactions with visitors.

[0532] "Speech recognition means" refers to a device or program that converts received speech data into text data.

[0533] A "natural language processing tool" is a program that analyzes text data, understands its meaning, and classifies it into appropriate categories.

[0534] A "response means" refers to a device or program for transmitting a generated response message to an visitor.

[0535] A "notification means" is a device or program that notifies the user's communication device of the generated response content or information regarding the visitor's purpose.

[0536] A "server system" is a computer system that activates generative artificial intelligence and controls and manages various systems in a centralized manner.

[0537] "Communication devices" refer to terminals that allow users to receive notifications, and include smartphones and tablets.

[0538] The system for implementing this invention automatically processes a series of actions from visitor detection to response and notification. The hardware includes sensors, terminals, servers, and communication equipment, while the software includes generative artificial intelligence, a speech recognition engine, a natural language processing engine, and a notification system.

[0539] First, the system uses sensors to detect visitors. For example, infrared sensors or motion sensors may be used. When a visitor is detected, the sensor generates an event signal, which is transmitted to the server via the network.

[0540] When the server receives this event signal, it activates a generative artificial intelligence (e.g., the GPT-3 model). The generative AI then generates prompt sentences to start the initial conversation. For example, it might generate questions such as "Who is this?", "What can I help you with?", or "Who in your family is this addressed to?". The generated prompt sentences are then converted into speech data. A text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) performs this conversion.

[0541] The generated audio data is played back through the terminal's speaker, conveying the question to the visitor. Next, the intercom's microphone records the visitor's response in real time and sends the audio data to the server. The server uses speech recognition (e.g., IBM Watson Speech to Text API) to convert the received audio data into text data.

[0542] The converted text data is analyzed using natural language processing tools (e.g., the spaCy library), and the visitor's request is classified into several categories. For example, possible categories include receiving packages when absent, refusing solicitations, responding during remote meetings, and handling children left alone at home.

[0543] Based on the classification results, the generative artificial intelligence generates an appropriate response message. For example, a message such as, "I am currently out, please leave your package in front of the door." This response message is sent to the terminal and conveyed to the visitor through the intercom speaker. The generated response content and the visitor's request are recorded on the server and also pushed to the user's communication device using a notification method (e.g., Firebase Cloud Messaging).

[0544] Users can check notifications on communication devices such as smartphones and generate a follow-up message if necessary. By sending this follow-up message to the server, it becomes possible to provide visitors with appropriate instructions again.

[0545] Specific example

[0546] Example 1: Leaving delivery unattended when the recipient is absent.

[0547] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[0548] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[0549] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver your package," and sends the voice data to the server.

[0550] 4. Server: Uses speech recognition to convert speech data into text data, which is then analyzed using natural language processing. This is classified as "delivery left unattended when the recipient is absent."

[0551] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0552] 6. Terminal: Receives a response message from the visitor via the speaker.

[0553] 7. Server: Notifies the user's communication device with the message, "A delivery service has arrived. The package has been left in front of your door."

[0554] Example 2: Dealing with persistent solicitation / sales tactics

[0555] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[0556] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[0557] 3. Terminal: The visitor responds, "I've come to introduce you to a new service," and sends the audio data to the server.

[0558] 4. Server: Converts speech data into text data using speech recognition technology and analyzes it using natural language processing technology. Classified as solicitation / sales.

[0559] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0560] 6. Terminal: Receives a response message from the visitor via the speaker.

[0561] This system allows users to significantly automate their intercom responses, improving the convenience and safety of their daily lives.

[0562] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0563] Step 1:

[0564] The terminal detects visitors using the intercom's sensor. The sensor consists of infrared sensors and motion sensors, and generates an event signal when it detects a visitor's movement. The input is the presence of a visitor, and the output is the transmission of the event signal to the server.

[0565] Step 2:

[0566] The server receives an event signal from a sensor. Upon receiving the event signal, the server activates a generative artificial intelligence (e.g., the GPT-3 model) to generate a prompt for initial interaction. The input is the event signal, and the output is the prompt. Specifically, the generative artificial intelligence generates a question such as "Who are you?"

[0567] Step 3:

[0568] The server uses a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) to convert the generated prompt text into speech data. The input is the prompt text, and the output is speech data. The TTS engine outputs the prompt text as synthesized speech.

[0569] Step 4:

[0570] The terminal receives audio data transmitted from the server and plays it back through the intercom speaker. The input is audio data, and the output is the audio message conveyed to the visitor. The speaker plays the audio data and conveys the question to the visitor.

[0571] Step 5:

[0572] The terminal uses the intercom's microphone to record the visitor's response in real time. The recorded audio data is immediately sent to the server. The input is the visitor's voice, and the output is the transmission of audio data to the server.

[0573] Step 6:

[0574] The server converts the received audio data into text data using speech recognition software (e.g., IBM Watson Speech to Text API). The input is audio data, and the output is text data. The speech recognition software analyzes the audio and outputs it as text data.

[0575] Step 7:

[0576] The server sends text data to a natural language processing system (e.g., the spaCy library), which analyzes the content and classifies the visitor's request. The input is text data, and the output is the classification result. The natural language processing system performs analysis and classification, assigning the request to a specific category (e.g., receiving a package when absent, refusing a solicitation, etc.).

[0577] Step 8:

[0578] The server uses generative artificial intelligence to generate an appropriate response message based on the classification result. The input is the classification result, and the output is the response message. The generative artificial intelligence then generates a prompt sentence, which is used as the response message.

[0579] Step 9:

[0580] The terminal receives a response message (audio data) from the server and transmits it to the visitor through the intercom speaker. The input is the response message, and the output is the audio transmitted to the visitor. The speaker plays the response message and responds to the visitor.

[0581] Step 10:

[0582] The server records the response and visitor's request and sends a push notification to the user's device using a notification method (e.g., Firebase Cloud Messaging). The input is the response and visitor's request, and the output is a push notification to the user's device. The notification method generates the message and sends it to the user's device.

[0583] Step 11:

[0584] The user can review the notification and, if necessary, generate a reply message and send it to the server. The input is the user's reply message, and the output is the message sent to the server. The user inputs and sends the message using a smartphone or similar device.

[0585] Step 12:

[0586] The server processes the received reply message and transmits it to the visitor through the terminal. The input is the reply message, and the output is the audio message to be delivered to the visitor. The server converts the reply message into audio data and plays it through the terminal's speaker.

[0587] The above outlines the specific processing steps of this system. This automates the entire process from visitor detection to response and notification, significantly reducing the burden on the user.

[0588] (Application Example 1)

[0589] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0590] In recent years, many homes and businesses have frequently had to deal with visitors, and prompt and appropriate responses are especially required when people are away or busy. However, current intercom systems require manual responses, which is inefficient and poses security risks. Furthermore, it is difficult to respond quickly to unexpected visitors such as solicitors or salespeople, resulting in wasted time and effort for users. Moreover, if these responses are not handled properly, important notifications may be delayed, which can negatively impact users' quality of life.

[0591] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0592] In this invention, the server includes detection means for detecting visitors, computation means for activating generative artificial intelligence to conduct initial dialogue, speech recognition means for converting speech to text, natural language processing means for analyzing text and classifying the visitor's request, response means for generating and responding with an automated response message, speech synthesis means for generating a voice response, and push notification means for sending visitor information to the user's terminal. This makes it possible to automate visitor interactions efficiently and securely. Furthermore, it enables appropriate responses in various situations such as handling deliveries when the user is absent and refusing solicitations, reducing the burden on the user and improving security.

[0593] A "visitor" is a person who visits a specific location, such as a home or office, and attempts to make contact using an intercom.

[0594] A "detection means" is a device or sensor used to detect the presence of a visitor, and it plays the first role in recognizing a visitor.

[0595] "Generative artificial intelligence" refers to artificial intelligence technologies that automatically perform intelligent tasks such as conversation, text generation, and speech recognition.

[0596] "Computation means" refers to a part of a server or computer system that receives and processes data from sensors and other devices.

[0597] "Voice recognition means" refers to a technology or device that converts voice input into text data, and is a means of recording the content of conversations with visitors as text.

[0598] "Natural language processing" refers to technologies that classify and understand specific intentions and requirements through the analysis of text data.

[0599] A "response means" is a technology or device for generating an appropriate response message based on classified requirements and communicating it to visitors.

[0600] A "speech synthesis engine" is a technology that converts text data into speech data and outputs it as speech.

[0601] "Push notification methods" refer to technologies for sending notifications to a user's device in real time.

[0602] This invention is a system that automatically processes a series of actions from visitor detection to response and notification. This system mainly consists of the following components: a detection means for detecting visitors, a computation means for activating generative artificial intelligence and conducting an initial dialogue, a speech recognition means for converting speech to text, a natural language processing means for analyzing text data and classifying the visitor's request, a response means for generating a response message and communicating it to the visitor, a push notification means for sending visitor information to the user's terminal, and a speech synthesis engine.

[0603] When the server receives a signal from the sensor, it activates generative artificial intelligence and begins an initial conversation with the visitor. For example, it might ask questions such as, "Who is this? What can I help you with? Who in your family is this for?" The initial conversation is generated by a speech synthesis engine and output through the speaker. The visitor's responses are captured by the intercom microphone and converted into text through speech recognition.

[0604] Next, the server analyzes the converted text using natural language processing to classify the visitor's request. Based on the classified request, a generative artificial intelligence generates an appropriate response message and transmits it to the visitor via a response system.

[0605] For example, if a visitor responds, "It's a delivery service. I've come to deliver your package," the server generates a response message saying, "You are currently out, please leave your package in front of the door," and outputs it using a speech synthesis engine.

[0606] Furthermore, the server uses push notifications to inform the user's device of visitor information and response content. This notification is delivered in real time using APIs such as LINE Notify API. Users receive notifications through their devices and can generate custom response messages as needed, then transmit them to the visitor again via the server.

[0607] The following scenarios are possible as specific examples.

[0608] When a visitor presses the intercom, the following notification is sent to the smartphone: "Who is it? What can I do for you? Who in the family is this for?". If the visitor replies, "It's a delivery. I've come to deliver a package," the system responds, "You are currently out, please leave the package in front of the door," and simultaneously sends a notification to the user's smartphone saying, "A delivery person has arrived. They have left the package in front of the door."

[0609] In this way, this invention automates visitor reception, significantly reducing the burden on users and improving security.

[0610] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0611] Step 1:

[0612] The sensor detects the visitor.

[0613] Specific operation: The detection device (sensor) detects when a visitor presses the intercom and sends an event signal to the server. The input data is information about the visitor's presence, and the output data is the event signal sent to the server.

[0614] Step 2:

[0615] The server receives an event signal and activates the generative artificial intelligence.

[0616] Specific operation: The server receives an event signal from the sensor and activates a generative artificial intelligence (AI). The generative AI prepares initial questions and uses a speech synthesis engine to output a voice message such as, "Who is this? What can I help you with? Who in your family is this for?" The input data is the event signal, and the output data is a voice message to the visitor.

[0617] Step 3:

[0618] The device captures the visitor's response and sends it to the server.

[0619] Specific operation: The intercom's microphone captures the visitor's response and sends it to the server as audio data. The input data is the visitor's voice, and the output data is the audio data sent to the server.

[0620] Step 4:

[0621] The server receives the audio data and converts it into text using speech recognition technology.

[0622] Specific operation: The server converts audio data into text data using speech recognition technology. The input data is audio data, and the output data is the converted text data.

[0623] Step 5:

[0624] The server analyzes text data using natural language processing to classify the visitor's purpose.

[0625] Specific operation: The server analyzes text data using natural language processing and classifies it into categories such as "delivery" and "sales." Input data is text data, and output data is classified category information.

[0626] Step 6:

[0627] The server uses generative artificial intelligence to generate response messages based on classified requests, and a speech synthesis engine creates voice messages.

[0628] Specific operation: The server uses generative artificial intelligence to generate appropriate response messages for classified requests and outputs them as voice messages using a speech synthesis engine. For example, if a visitor says "I have a delivery," the server will generate a message saying, "I am currently out, please leave your package in front of the door." The input data is classified request information, and the output data is the response voice message.

[0629] Step 7:

[0630] The terminal transmits the generated response message to the visitor.

[0631] Specific operation: The generated response message is played to the visitor through the device's speaker. The input data is the response voice message, and the output data is the voice output to the visitor.

[0632] Step 8:

[0633] The server creates and sends a push notification to the user's device containing visitor information and response details.

[0634] Specific operation: The server uses push notification methods, such as the LINE Notify API, to send visitor information and response content as push notifications to the user's smartphone. The input data consists of the response content and visitor information, and the output data is the push notification message sent to the user's device.

[0635] Step 9:

[0636] The user generates a follow-up message as needed and sends it to the server.

[0637] Specific operation: The user receives a push notification on their smartphone, enters a reply message as needed, and sends it to the server. The input data is the user's reply message, and the output data is the reply data sent to the server.

[0638] Step 10:

[0639] The server receives a response message from the user, converts it into an appropriate format using generative artificial intelligence, and transmits it to the visitor via the terminal.

[0640] Specific operation: The server receives a response message from the user, converts it into a voice message using generative artificial intelligence, and plays it back to the visitor through the terminal's speaker. The input data is the user's response message, and the output data is the voice message to the visitor.

[0641] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0642] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[0643] Initial processing when the intercom rings

[0644] Terminal:

[0645] When a visitor arrives at the intercom, a sensor detects this and sends an event signal to the server. This allows the server to recognize the visitor's presence.

[0646] server:

[0647] Upon receiving the event signal, the server activates the generative artificial intelligence and initiates an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[0648] Dialogue with visitors

[0649] Terminal:

[0650] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[0651] server:

[0652] The received audio data is converted into text data using speech recognition technology. This text data accurately reflects the content of the visitor's speech.

[0653] server:

[0654] The converted text data is analyzed using natural language processing to classify the visitor's purpose into multiple categories. Examples of these categories include: unattended delivery, solicitation, remote meeting, and childcare. Based on the analysis results, an appropriate automated response message is generated by a generative artificial intelligence system.

[0655] Automated responses and user notifications

[0656] server:

[0657] The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. This ensures that the visitor receives an appropriate response.

[0658] server:

[0659] The generated response and information regarding the visitor's purpose are sent to the user's device using a notification system.

[0660] User emotion recognition and adjustment

[0661] server:

[0662] An emotion engine built into the user's device recognizes emotions from the user's text or voice input. Based on this emotional state, notification content can be adjusted. For example, if the user is stressed, notifications will be delivered using softer language.

[0663] server:

[0664] Furthermore, the emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[0665] User response

[0666] User:

[0667] Users who receive notifications on devices such as smartphones can generate additional instructions or response messages via the server as needed. For example, if a friend comes to visit, they can send a custom message such as, "I'll be right back, please wait a moment."

[0668] Specific example

[0669] Example 1: Leaving delivery unattended when the recipient is absent.

[0670] 1. Terminal: The intercom rings.

[0671] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0672] 3. Terminal: The visitor responds, "Delivery. I've come to deliver your package."

[0673] 4. Server: Receives voice data and converts it to text. Analyzes it using natural language processing to determine whether to leave the package unattended when the recipient is absent.

[0674] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0675] 6. Terminal: Provides instructions to visitors.

[0676] 7. Server: Notifies the user that "Your delivery has arrived. The package has been left in front of your door."

[0677] 8. Server: If the emotion engine detects the user's stress level, it adjusts the notification content to a softer tone.

[0678] Example 2: Solicitation / Sales Strategies

[0679] 1. Terminal: The intercom rings.

[0680] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0681] 3. Terminal: The visitor responds, "This is to inform you about a new service."

[0682] 4. Server: Receives audio data, converts it to text, analyzes it using natural language processing, and determines whether it is a solicitation / sales call.

[0683] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0684] 6. Terminal: Provides instructions to visitors.

[0685] 7. Server: Notifies the user that "a solicitation was received and declined."

[0686] 8. Server: The emotion engine dynamically adjusts the tone and content of notifications according to the user's situation.

[0687] In this way, it is possible to efficiently handle interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[0688] The following describes the processing flow.

[0689] Step 1:

[0690] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[0691] Step 2:

[0692] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[0693] Step 3:

[0694] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[0695] Step 4:

[0696] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[0697] Step 5:

[0698] Terminal: Sends recorded audio data to the server via the internet.

[0699] Step 6:

[0700] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[0701] Step 7:

[0702] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[0703] Step 8:

[0704] Server: Based on the analysis results, the AI ​​generates an appropriate automated response message. This message includes the most suitable response for the request.

[0705] Step 9:

[0706] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[0707] Step 10:

[0708] Terminal: The response message is properly conveyed to the visitor.

[0709] Step 11:

[0710] Server: Sends the generated response and information about the visitor's purpose to the user's terminal using a notification system.

[0711] Step 12:

[0712] User: The user receives notifications on a device such as a smartphone, and the emotion engine analyzes the user's emotions.

[0713] Step 13:

[0714] Server: The emotion engine recognizes the user's emotional state and adjusts the notification content as needed. For example, if the user is feeling stressed, a notification will be sent in softer language.

[0715] Step 14:

[0716] Server: The emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[0717] Step 15:

[0718] User: Check the notification, generate a reply message from your smartphone if necessary, and transmit it to the visitor via the server. For example, you can send a custom message such as, "I'll be right back, please wait a moment."

[0719] These steps enable the system to efficiently automate visitor interactions while considering the user's emotional state, thereby improving user safety and convenience.

[0720] (Example 2)

[0721] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0722] Traditional intercom systems require manual responses to visitors, making it difficult to provide appropriate care when the user is absent or busy. Furthermore, they lack the ability to respond flexibly to emotions and situations, hindering user convenience and safety. Additionally, they often fail to accurately identify the visitor's purpose and situation, leading to delays in responses.

[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0724] In this invention, the server includes a sensor means for detecting visitors, a computing means for activating generative artificial intelligence to conduct initial dialogue, and a speech recognition means for converting speech to text. This enables automatic detection of visitors and appropriate initial dialogue. It also includes a natural language processing means for analyzing text and classifying the visitor's request, and an emotion engine for recognizing the user's emotions and adjusting notification content and responses. This enables flexible responses that take into account the user's emotional state, improving user safety and convenience.

[0725] A "visitor" is a person who visits via the intercom system.

[0726] "Sensor means" refers to devices or means used to detect the presence of visitors, and includes motion detection and infrared sensors.

[0727] "Computer means" refers to a system or device that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[0728] "Generative artificial intelligence" refers to artificial intelligence technology that generates voice dialogues and response messages within a system.

[0729] "Voice recognition means" refers to technologies and devices for converting a visitor's voice into text data.

[0730] "Natural language processing means" refers to technologies and devices that analyze text data and classify the purpose of a visitor's visit.

[0731] A "response means" refers to a device or technology that generates an automated response message according to the classified request and transmits it to the visitor.

[0732] "Notification means" refers to technologies and devices used to notify the user's terminal of the content of the response or the purpose of the visitor's visit.

[0733] An "emotion engine" is a technology or system that recognizes a user's emotions and adjusts notification content and responses accordingly.

[0734] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, computer means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[0735] Initial processing when the intercom rings

[0736] When the terminal detects a visitor, the sensor detects this and sends an event signal to the server. Upon receiving the event signal, the server activates generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who are you?", "What can I help you with?", and "Who in your family is this for?". The generative artificial intelligence uses speech synthesis technology to output this as speech.

[0737] Dialogue with visitors

[0738] The device's microphone captures the visitor's responses and records them in real time. The recorded audio data is sent to a server. The server converts the received audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). After conversion, the text data is analyzed by a natural language processing unit (e.g., Hugging Face's NLP model) to classify the visitor's request into multiple categories.

[0739] Automated responses and user notifications

[0740] The server generates an appropriate automated response message based on the analysis results. The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. In addition, information regarding the response content and the visitor's purpose is sent to the user's terminal using notification methods (e.g., push notification to a smartphone or email).

[0741] User emotion recognition and adjustment

[0742] The server has an emotion engine that can recognize emotions from text and voice input from the user's device. The emotion engine utilizes the Emotion API and other tools to adjust notification content to a softer tone if the user is feeling stressed. The emotion engine can also dynamically adjust the tone and content of automated response messages based on the user's emotional state. If the user is relaxed, messages will be generated in a friendly tone.

[0743] User response

[0744] When a user receives a notification on a device such as a smartphone, they can generate additional instructions or reply messages via the server as needed. For example, if a friend visits, they can send a custom message such as, "I'll be right back, please wait a moment."

[0745] Specific example

[0746] Example 1: Leaving delivery unattended when the recipient is absent.

[0747] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[0748] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[0749] 3. Terminal: The intercom's microphone captures the visitor's speech, such as "Delivery. I've come to deliver your package," and sends it to the server.

[0750] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[0751] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if the package was left unattended when the recipient was absent.

[0752] 6. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0753] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[0754] 8. Server: Notifies the user, "The delivery has arrived. The package has been left in front of your door."

[0755] 9. Server: The server uses the Emotion API to sense the user's stress level and adjusts the notification content to a softer tone.

[0756] Example of a prompt:

[0757] "If your doorbell rings and a visitor asks you to leave a package, how would you respond?"

[0758] Example 2: Solicitation / Sales Strategies

[0759] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[0760] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[0761] 3. Terminal: The intercom's microphone captures the visitor's statement, "We have an announcement about a new service," and sends it to the server.

[0762] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[0763] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if it is a solicitation / sales pitch.

[0764] 6. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0765] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[0766] 8. Server: Notifies the user that "a solicitation was received and declined."

[0767] 9. Server: Use the Emotion API to check the user's status and dynamically adjust the tone and content of notifications.

[0768] In this way, the system efficiently handles interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[0769] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0770] Step 1:

[0771] Terminal: The intercom's sensor detects a visitor. As input, the sensor obtains information about the visitor's presence. When the sensor is triggered, an event signal is generated and sent to the server. As output, the server recognizes the visitor's presence.

[0772] Step 2:

[0773] Server: The server receives an event signal and activates the generative artificial intelligence. It receives the event signal as input. The generative artificial intelligence generates initial dialogue questions. It generates questions such as "Who is this?", "What can I help you with?", and "Who in your family is this addressed to?" as text and converts them into speech using speech synthesis technology. It generates the voiced questions as output and sends them to the terminal.

[0774] Step 3:

[0775] Terminal: The intercom's microphone captures and records the visitor's response. It receives the visitor's voice response as input. The recorded data is sent to the server. It also sends the voice data to the server as output.

[0776] Step 4:

[0777] Server: The server converts the received audio data into text data using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). It receives audio data as input. The speech recognition engine analyzes the audio and generates text data. It generates text data as output.

[0778] Step 5:

[0779] Server: The converted text data is analyzed using natural language processing (e.g., Hugging Face's NLP model). It receives text data as input. The natural language processing analyzes the text content and classifies the visitor's request into multiple categories. It generates classified category information as output.

[0780] Step 6:

[0781] Server: Generates an appropriate automated response message based on the analysis results. It receives classified category information as input. The generative artificial intelligence generates a response message such as, "We are currently out, please leave your package at the front door." It sends the generated response message to the terminal as output.

[0782] Step 7:

[0783] Terminal: The terminal transmits the response message received from the server to the visitor through the intercom speaker. Input: Receives the response message. Plays the voice message using the intercom speaker. Output: Transmits the response message to the visitor.

[0784] Step 8:

[0785] Server: The generated response and information regarding the visitor's request are sent to the user's device using a notification method. The server receives the response and request information as input. It then notifies the user's device using a notification method (e.g., push notification or email). The notification is then delivered to the user as output.

[0786] Step 9:

[0787] Server: The server contains an emotion engine that recognizes the user's emotions from text and voice input from the user's device. It receives user input data as input. It analyzes the user's emotional state using the Emotion API and other tools. It generates user emotion information as output.

[0788] Step 10:

[0789] Server: The emotion engine adjusts notification content based on the user's emotional state. It receives user emotion information as input. It adjusts the notification content and tone according to the emotional state to generate an appropriate notification message. As output, it sends the adjusted notification message to the user's device.

[0790] Step 11:

[0791] User: The user receives a notification on a device such as a smartphone and generates additional instructions or reply messages via the server as needed. Input: Receives the user notification. Creates a custom message (e.g., "I'll be right back, please wait a moment") as needed and sends it to the server. Output: The reply message is transmitted to the visitor.

[0792] (Application Example 2)

[0793] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0794] Traditional intercom systems often required manual interaction with visitors, making quick and appropriate responses difficult, especially when users were away from home, working remotely, or with children home alone. Furthermore, the lack of appropriate notifications and responses tailored to the user's emotional state often led to inconvenience and stress. As a result, many users felt a lack of both safety and convenience.

[0795] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to perform an initial dialogue, a means for the generative artificial intelligence to convert speech into text, a means for analyzing the text and classifying the visitor's request, a means for generating an automatic response message according to the classified request and transmitting it to the visitor, a means for notifying the request and response content to the user's communication terminal, and a means for recognizing the user's emotions and adjusting the tone and content of the notification and automatic response message. This enables quick and appropriate responses even when the user is absent or busy, and allows for flexible notifications and responses according to the user's emotional state.

[0796] "Sensor means" refers to a device or system used to detect visitors.

[0797] A "server means" is a device or system that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[0798] "Generative artificial intelligence" refers to artificial intelligence software that converts speech into text, analyzes that text, and generates appropriate response messages.

[0799] "Voice recognition means" refers to a device or system that converts a visitor's voice into text.

[0800] A "natural language processing device" is a device or system that analyzes text to classify the visitor's request.

[0801] "Response means" refers to a device or system that generates an automated response message according to the classified requirements and transmits it to the visitor.

[0802] "Notification means" refers to a device or system that notifies the user's communication terminal of the aforementioned request and response content.

[0803] "Emotion recognition means" refers to a device or system that recognizes the user's emotions and adjusts the tone and content of the notification and automated response messages.

[0804] A "communication terminal" is an electronic device, including smartphones and tablets, that a user uses to receive notifications.

[0805] The system for implementing this invention automatically performs a series of operations from visitor detection to response and notification, and further recognizes the user's emotions to adjust notifications and responses accordingly. The specific form of this system is described below.

[0806] Hardware configuration

[0807] This system uses the following hardware:

[0808] Sensory devices (e.g., intercoms, security cameras, door sensors)

[0809] Server configuration (e.g., cloud server)

[0810] User's communication device (e.g., smartphone, tablet)

[0811] Software to use

[0812] This system is equipped with the following software:

[0813] Generative artificial intelligence (e.g. GPT-3)

[0814] Speech recognition methods (e.g., Google Speech Recognition API)

[0815] Natural language processing tools (e.g., TextBlob library)

[0816] Emotion recognition methods (e.g., TextBlob's sentiment analysis function)

[0817] Data processing and data calculation

[0818] The server performs the following data processing and calculations.

[0819] 1. Visitor detection

[0820] Sensors detect the presence of visitors. Examples include intercoms and door sensors.

[0821] 2. Initiating the initial dialogue

[0822] The server receives a notification from a sensor and activates a generative artificial intelligence (AI) to initiate an initial conversation with the visitor, such as "Who are you?". This AI then uses speech recognition to convert the visitor's voice into text.

[0823] 3. Classification of Visitor's Purpose

[0824] Text data converted by generative artificial intelligence is analyzed using natural language processing techniques to appropriately classify the visitor's purpose.

[0825] 4. Generation and transmission of response messages

[0826] Based on the categorized request, an automated response message is generated and communicated to the visitor via the response system.

[0827] 5. Notification to the user's terminal

[0828] Simultaneously, the server sends information about the response and the visitor's purpose to the user's communication terminal using a notification mechanism.

[0829] 6. User emotion recognition and notification adjustment

[0830] Furthermore, using emotion recognition mechanisms, the system recognizes the user's emotions from their text or voice input and dynamically adjusts the tone of notification content and response messages.

[0831] Specific example

[0832] The following are specific examples.

[0833] Example 1: Delivery company's response

[0834] 1. The sensor detects the delivery person.

[0835] 2. The server activates a generative artificial intelligence and responds, "Who is this?"

[0836] 3. The delivery person responds, "I've come to deliver your package."

[0837] 4. The speech recognition means converts the speech into text, and the natural language processing means analyzes the text.

[0838] 5. The server generates a response message saying, "We are currently out, please leave it at the front door," and transmits it to the delivery person via the response system.

[0839] 6. The user's communication device is notified with the message, "Your delivery has arrived. The package has been left in front of your door."

[0840] Example 2: Detection of a suspicious person

[0841] 1. The sensor detects a suspicious person.

[0842] 2. The server activates a generative artificial intelligence and responds, "Who is it?", but the suspicious person remains silent.

[0843] 3. The server identifies the person as suspicious, generates a warning message, and transmits it via the response system.

[0844] 4. Simultaneously, a warning notification is sent to the user's communication terminal.

[0845] Example of a prompt

[0846] "Person says: I'm here to deliver a package. User is stressed. Generate a soft response."

[0847] In this way, it becomes possible to respond quickly and appropriately even when the user is away or busy, and to provide flexible notifications and responses that are tailored to their emotional state.

[0848] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0849] Step 1:

[0850] A sensor detects the presence of a visitor. The input is the sensor (e.g., intercom, security camera, door sensor), and the output is a notification signal of the detection event. This notification signal is sent to the server.

[0851] Step 2:

[0852] The server receives a notification from a sensor. The input is a notification signal, and the server activates a generative artificial intelligence to initiate an initial conversation. Specifically, it generates questions such as "Who are you?" for the visitor and outputs them as voice through a response system.

[0853] Step 3:

[0854] The visitor's response is captured by the intercom's microphone and recorded in real time as audio data. The input is the visitor's voice, and the output is the recorded audio data. The terminal sends this audio data to the server.

[0855] Step 4:

[0856] The server converts the received audio data into text data using speech recognition tools (e.g., Google Speech Recognition API). The input is audio data, and the output is text data. Specifically, the speech recognition engine analyzes the audio waveform data and converts it into the corresponding text.

[0857] Step 5:

[0858] The server analyzes the converted text data using natural language processing tools (e.g., the TextBlob library). The input is text data, and the output is the category of requests resulting from the analysis. In this process, the content of the text is analyzed, and its intent and requirements are classified into categories such as receiving packages when absent, refusing solicitations, responding during remote meetings, and how to handle situations when children are home alone.

[0859] Step 6:

[0860] The server generates an automated response message based on the analysis results. The input is the category of the request, and the output is the automated response message. Specifically, a generative artificial intelligence generates an appropriate message and conveys it to the visitor. For example, if the recipient is absent and the package is to be left at the door, the server will generate a message such as, "We are currently absent, please leave your package at the front door."

[0861] Step 7:

[0862] The server transmits a generated response message to the visitor via the response mechanism. The input is an automated response message, and the output is an audio message to the visitor. Specifically, it is transmitted to the visitor through the intercom speaker.

[0863] Step 8:

[0864] The server sends the generated response and information about the visitor's purpose to the user's communication terminal using a notification mechanism. The input is the response and purpose information, and the output is a notification message to the user's terminal. Specifically, it is displayed as a notification on the user's smartphone.

[0865] Step 9:

[0866] The server uses emotion recognition to recognize the user's emotions and adjusts the tone and content of notifications and automated response messages accordingly. The input is the user's emotion data (e.g., text or voice input), and the output is the adjusted notification and response message. For example, if the user is feeling stressed, the notification content might be adjusted to a softer tone, such as "Your package has been left at your front door."

[0867] Step 10:

[0868] The user generates additional instructions or follow-up messages as needed, which are then transmitted to the visitor via the server. The input is the user's follow-up message, and the output is a voice message to the visitor. For example, the user might send a custom message such as, "I'll be right back, please wait a moment."

[0869] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0870] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0871] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0872] [Third Embodiment]

[0873] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0874] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0875] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0876] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0877] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0878] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0879] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0880] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0881] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0882] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0883] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0884] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0885] The system for carrying out this invention automatically processes a series of operations from visitor detection to response and notification. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, and notification means.

[0886] Initial processing when the intercom rings

[0887] Terminal:

[0888] When a visitor arrives at the intercom, a sensor detects them and sends an event signal to the server. This allows the server to recognize that a visitor has arrived.

[0889] server:

[0890] Upon receiving the event signal, the server activates the generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[0891] Dialogue with visitors

[0892] Terminal:

[0893] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[0894] server:

[0895] The server converts the received audio data into text data using speech recognition. The converted text data is then analyzed by natural language processing to classify the visitor's request.

[0896] Classification of inquiries and automated responses

[0897] server:

[0898] Natural language processing tools analyze text data and classify visitors' requests into several categories. These categories include, for example, unattended delivery, solicitation or sales calls, remote meetings, and children home alone. An automated response message tailored to each request is then generated by a generative artificial intelligence system.

[0899] Terminal:

[0900] The generated response message is transmitted to the visitor through the device's speaker, ensuring they receive an appropriate response.

[0901] Notification to the user

[0902] server:

[0903] The generated response and information regarding the visitor's purpose are sent as push notifications to the user's device using a notification system. For example, if the user has a smartphone, a specific notification such as "We have confirmed that Sagawa Express has arrived and left the package at your front door" will be sent.

[0904] User:

[0905] Users can review received notifications, generate a follow-up message if necessary, and transmit it to visitors via the server. This follow-up allows users to send custom messages, such as "I'll be right back, please wait a moment," when a friend visits.

[0906] Specific example

[0907] Here are some specific examples.

[0908] Example 1: Leaving delivery unattended when the recipient is absent.

[0909] 1. Terminal: The intercom rings.

[0910] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0911] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver a package."

[0912] 4. Server: Receives voice data, converts it to text, and analyzes it using natural language processing. Determines that it is a delivery attempt when the recipient is absent.

[0913] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0914] 6. Terminal: Provides instructions to visitors.

[0915] 7. Server: Notifies the user that "A delivery service has arrived. The package has been left in front of your door."

[0916] Example 2: Dealing with persistent solicitation / sales tactics

[0917] 1. Terminal: The intercom rings.

[0918] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[0919] 3. Terminal: The visitor responds, "I've come to introduce a new service."

[0920] 4. Server: Receives audio data and converts it to text. Analysis using natural language processing techniques indicates that it is a solicitation / sales pitch.

[0921] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0922] 6. Terminal: Provides instructions to visitors.

[0923] This system can streamline intercom communication in users' daily lives, improving security and convenience.

[0924] The following describes the processing flow.

[0925] Step 1:

[0926] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[0927] Step 2:

[0928] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[0929] Step 3:

[0930] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[0931] Step 4:

[0932] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[0933] Step 5:

[0934] Terminal: Sends recorded audio data to the server via the internet.

[0935] Step 6:

[0936] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[0937] Step 7:

[0938] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[0939] Step 8:

[0940] Server: Based on the classification results, an appropriate automated response message is generated by the AI ​​generation system. This message includes the optimal response method for the given request.

[0941] Step 9:

[0942] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[0943] Step 10:

[0944] Terminal: The response message is properly conveyed to the visitor, and the response ends.

[0945] Step 11:

[0946] Server: Notifies the user's terminal of the generated response and information regarding the visitor's purpose.

[0947] Step 12:

[0948] Users can receive notifications on devices such as smartphones, and, if necessary, generate additional instructions or response messages via the server and communicate them to visitors.

[0949] These steps automate the intercom system, allowing for efficient handling of visitors without direct user interaction.

[0950] (Example 1)

[0951] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0952] Traditional intercom systems require users to manually handle a series of processes, from detecting visitors to answering, providing appropriate responses, and notifying the user, which places a significant burden on the user. This inconvenience is particularly noticeable when users are away for extended periods or working remotely. Furthermore, dealing with inappropriate solicitations and sales calls presents challenges, highlighting the need for automated systems to address these issues.

[0953] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0954] In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to conduct an initial dialogue, a speech recognition means in which the generative artificial intelligence converts speech into text, a natural language processing means for analyzing the text and classifying the visitor's request, a response means for generating an automated response message according to the classified request and communicating it to the visitor, and a notification means for notifying the user's communication device of the request and response content. As a result, a series of processes from visitor detection to response and notification are automated, significantly reducing the burden on the user, improving convenience, and ensuring security.

[0955] A "sensor device" is a device that detects visitors, generates an event signal, and notifies the server.

[0956] "Generative artificial intelligence" is a program that generates text questions and answers to automate interactions with visitors.

[0957] "Speech recognition means" refers to a device or program that converts received speech data into text data.

[0958] A "natural language processing tool" is a program that analyzes text data, understands its meaning, and classifies it into appropriate categories.

[0959] A "response means" refers to a device or program for transmitting a generated response message to an visitor.

[0960] A "notification means" is a device or program that notifies the user's communication device of the generated response content or information regarding the visitor's purpose.

[0961] A "server system" is a computer system that activates generative artificial intelligence and controls and manages various systems in a centralized manner.

[0962] "Communication devices" refer to terminals that allow users to receive notifications, and include smartphones and tablets.

[0963] The system for implementing this invention automatically processes a series of actions from visitor detection to response and notification. The hardware includes sensors, terminals, servers, and communication equipment, while the software includes generative artificial intelligence, a speech recognition engine, a natural language processing engine, and a notification system.

[0964] First, the system uses sensors to detect visitors. For example, infrared sensors or motion sensors may be used. When a visitor is detected, the sensor generates an event signal, which is transmitted to the server via the network.

[0965] When the server receives this event signal, it activates a generative artificial intelligence (e.g., the GPT-3 model). The generative AI then generates prompt sentences to start the initial conversation. For example, it might generate questions such as "Who is this?", "What can I help you with?", or "Who in your family is this addressed to?". The generated prompt sentences are then converted into speech data. A text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) performs this conversion.

[0966] The generated audio data is played back through the terminal's speaker, conveying the question to the visitor. Next, the intercom's microphone records the visitor's response in real time and sends the audio data to the server. The server uses speech recognition (e.g., IBM Watson Speech to Text API) to convert the received audio data into text data.

[0967] The converted text data is analyzed using natural language processing tools (e.g., the spaCy library), and the visitor's request is classified into several categories. For example, possible categories include receiving packages when absent, refusing solicitations, responding during remote meetings, and handling children left alone at home.

[0968] Based on the classification results, the generative artificial intelligence generates an appropriate response message. For example, a message such as, "I am currently out, please leave your package in front of the door." This response message is sent to the terminal and conveyed to the visitor through the intercom speaker. The generated response content and the visitor's request are recorded on the server and also pushed to the user's communication device using a notification method (e.g., Firebase Cloud Messaging).

[0969] Users can check notifications on communication devices such as smartphones and generate a follow-up message if necessary. By sending this follow-up message to the server, it becomes possible to provide visitors with appropriate instructions again.

[0970] Specific example

[0971] Example 1: Leaving delivery unattended when the recipient is absent.

[0972] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[0973] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[0974] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver your package," and sends the voice data to the server.

[0975] 4. Server: Uses speech recognition to convert speech data into text data, which is then analyzed using natural language processing. This is classified as "delivery left unattended when the recipient is absent."

[0976] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[0977] 6. Terminal: Receives a response message from the visitor via the speaker.

[0978] 7. Server: Notifies the user's communication device with the message, "A delivery service has arrived. The package has been left in front of your door."

[0979] Example 2: Dealing with persistent solicitation / sales tactics

[0980] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[0981] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[0982] 3. Terminal: The visitor responds, "I've come to introduce you to a new service," and sends the audio data to the server.

[0983] 4. Server: Converts speech data into text data using speech recognition technology and analyzes it using natural language processing technology. Classified as solicitation / sales.

[0984] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[0985] 6. Terminal: Receives a response message from the visitor via the speaker.

[0986] This system allows users to significantly automate their intercom responses, improving the convenience and safety of their daily lives.

[0987] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0988] Step 1:

[0989] The terminal detects visitors using the intercom's sensor. The sensor consists of infrared sensors and motion sensors, and generates an event signal when it detects a visitor's movement. The input is the presence of a visitor, and the output is the transmission of the event signal to the server.

[0990] Step 2:

[0991] The server receives an event signal from a sensor. Upon receiving the event signal, the server activates a generative artificial intelligence (e.g., the GPT-3 model) to generate a prompt for initial interaction. The input is the event signal, and the output is the prompt. Specifically, the generative artificial intelligence generates a question such as "Who are you?"

[0992] Step 3:

[0993] The server uses a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) to convert the generated prompt text into speech data. The input is the prompt text, and the output is speech data. The TTS engine outputs the prompt text as synthesized speech.

[0994] Step 4:

[0995] The terminal receives audio data transmitted from the server and plays it back through the intercom speaker. The input is audio data, and the output is the audio message conveyed to the visitor. The speaker plays the audio data and conveys the question to the visitor.

[0996] Step 5:

[0997] The terminal uses the intercom's microphone to record the visitor's response in real time. The recorded audio data is immediately sent to the server. The input is the visitor's voice, and the output is the transmission of audio data to the server.

[0998] Step 6:

[0999] The server converts the received audio data into text data using speech recognition software (e.g., IBM Watson Speech to Text API). The input is audio data, and the output is text data. The speech recognition software analyzes the audio and outputs it as text data.

[1000] Step 7:

[1001] The server sends text data to a natural language processing system (e.g., the spaCy library), which analyzes the content and classifies the visitor's request. The input is text data, and the output is the classification result. The natural language processing system performs analysis and classification, assigning the request to a specific category (e.g., receiving a package when absent, refusing a solicitation, etc.).

[1002] Step 8:

[1003] The server uses generative artificial intelligence to generate an appropriate response message based on the classification result. The input is the classification result, and the output is the response message. The generative artificial intelligence then generates a prompt sentence, which is used as the response message.

[1004] Step 9:

[1005] The terminal receives a response message (audio data) from the server and transmits it to the visitor through the intercom speaker. The input is the response message, and the output is the audio transmitted to the visitor. The speaker plays the response message and responds to the visitor.

[1006] Step 10:

[1007] The server records the response and visitor's request and sends a push notification to the user's device using a notification method (e.g., Firebase Cloud Messaging). The input is the response and visitor's request, and the output is a push notification to the user's device. The notification method generates the message and sends it to the user's device.

[1008] Step 11:

[1009] The user can review the notification and, if necessary, generate a reply message and send it to the server. The input is the user's reply message, and the output is the message sent to the server. The user inputs and sends the message using a smartphone or similar device.

[1010] Step 12:

[1011] The server processes the received reply message and transmits it to the visitor through the terminal. The input is the reply message, and the output is the audio message to be delivered to the visitor. The server converts the reply message into audio data and plays it through the terminal's speaker.

[1012] The above outlines the specific processing steps of this system. This automates the entire process from visitor detection to response and notification, significantly reducing the burden on the user.

[1013] (Application Example 1)

[1014] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1015] In recent years, many homes and businesses have frequently had to deal with visitors, and prompt and appropriate responses are especially required when people are away or busy. However, current intercom systems require manual responses, which is inefficient and poses security risks. Furthermore, it is difficult to respond quickly to unexpected visitors such as solicitors or salespeople, resulting in wasted time and effort for users. Moreover, if these responses are not handled properly, important notifications may be delayed, which can negatively impact users' quality of life.

[1016] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1017] In this invention, the server includes detection means for detecting visitors, computation means for activating generative artificial intelligence to conduct initial dialogue, speech recognition means for converting speech to text, natural language processing means for analyzing text and classifying the visitor's request, response means for generating and responding with an automated response message, speech synthesis means for generating a voice response, and push notification means for sending visitor information to the user's terminal. This makes it possible to automate visitor interactions efficiently and securely. Furthermore, it enables appropriate responses in various situations such as handling deliveries when the user is absent and refusing solicitations, reducing the burden on the user and improving security.

[1018] A "visitor" is a person who visits a specific location, such as a home or office, and attempts to make contact using an intercom.

[1019] A "detection means" is a device or sensor used to detect the presence of a visitor, and it plays the first role in recognizing a visitor.

[1020] "Generative artificial intelligence" refers to artificial intelligence technologies that automatically perform intelligent tasks such as conversation, text generation, and speech recognition.

[1021] "Computation means" refers to a part of a server or computer system that receives and processes data from sensors and other devices.

[1022] "Voice recognition means" refers to a technology or device that converts voice input into text data, and is a means of recording the content of conversations with visitors as text.

[1023] "Natural language processing" refers to technologies that classify and understand specific intentions and requirements through the analysis of text data.

[1024] A "response means" is a technology or device for generating an appropriate response message based on classified requirements and communicating it to visitors.

[1025] A "speech synthesis engine" is a technology that converts text data into speech data and outputs it as speech.

[1026] "Push notification methods" refer to technologies for sending notifications to a user's device in real time.

[1027] This invention is a system that automatically processes a series of actions from visitor detection to response and notification. This system mainly consists of the following components: a detection means for detecting visitors, a computation means for activating generative artificial intelligence and conducting an initial dialogue, a speech recognition means for converting speech to text, a natural language processing means for analyzing text data and classifying the visitor's request, a response means for generating a response message and communicating it to the visitor, a push notification means for sending visitor information to the user's terminal, and a speech synthesis engine.

[1028] When the server receives a signal from the sensor, it activates generative artificial intelligence and begins an initial conversation with the visitor. For example, it might ask questions such as, "Who is this? What can I help you with? Who in your family is this for?" The initial conversation is generated by a speech synthesis engine and output through the speaker. The visitor's responses are captured by the intercom microphone and converted into text through speech recognition.

[1029] Next, the server analyzes the converted text using natural language processing to classify the visitor's request. Based on the classified request, a generative artificial intelligence generates an appropriate response message and transmits it to the visitor via a response system.

[1030] For example, if a visitor responds, "It's a delivery service. I've come to deliver your package," the server generates a response message saying, "You are currently out, please leave your package in front of the door," and outputs it using a speech synthesis engine.

[1031] Furthermore, the server uses push notifications to inform the user's device of visitor information and response content. This notification is delivered in real time using APIs such as LINE Notify API. Users receive notifications through their devices and can generate custom response messages as needed, then transmit them to the visitor again via the server.

[1032] The following scenarios are possible as specific examples.

[1033] When a visitor presses the intercom, the following notification is sent to the smartphone: "Who is it? What can I do for you? Who in the family is this for?". If the visitor replies, "It's a delivery. I've come to deliver a package," the system responds, "You are currently out, please leave the package in front of the door," and simultaneously sends a notification to the user's smartphone saying, "A delivery person has arrived. They have left the package in front of the door."

[1034] In this way, this invention automates visitor reception, significantly reducing the burden on users and improving security.

[1035] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1036] Step 1:

[1037] The sensor detects the visitor.

[1038] Specific operation: The detection device (sensor) detects when a visitor presses the intercom and sends an event signal to the server. The input data is information about the visitor's presence, and the output data is the event signal sent to the server.

[1039] Step 2:

[1040] The server receives an event signal and activates the generative artificial intelligence.

[1041] Specific operation: The server receives an event signal from the sensor and activates a generative artificial intelligence (AI). The generative AI prepares initial questions and uses a speech synthesis engine to output a voice message such as, "Who is this? What can I help you with? Who in your family is this for?" The input data is the event signal, and the output data is a voice message to the visitor.

[1042] Step 3:

[1043] The device captures the visitor's response and sends it to the server.

[1044] Specific operation: The intercom's microphone captures the visitor's response and sends it to the server as audio data. The input data is the visitor's voice, and the output data is the audio data sent to the server.

[1045] Step 4:

[1046] The server receives the audio data and converts it into text using speech recognition technology.

[1047] Specific operation: The server converts audio data into text data using speech recognition technology. The input data is audio data, and the output data is the converted text data.

[1048] Step 5:

[1049] The server analyzes text data using natural language processing to classify the visitor's purpose.

[1050] Specific operation: The server analyzes text data using natural language processing and classifies it into categories such as "delivery" and "sales." Input data is text data, and output data is classified category information.

[1051] Step 6:

[1052] The server uses generative artificial intelligence to generate response messages based on classified requests, and a speech synthesis engine creates voice messages.

[1053] Specific operation: The server uses generative artificial intelligence to generate appropriate response messages for classified requests and outputs them as voice messages using a speech synthesis engine. For example, if a visitor says "I have a delivery," the server will generate a message such as "I am currently out, please leave your package in front of the door." The input data is classified request information, and the output data is the response voice message.

[1054] Step 7:

[1055] The terminal transmits the generated response message to the visitor.

[1056] Specific operation: The generated response message is played to the visitor through the device's speaker. The input data is the response voice message, and the output data is the voice output to the visitor.

[1057] Step 8:

[1058] The server creates and sends a push notification to the user's device containing visitor information and response details.

[1059] Specific operation: The server uses push notification methods, such as the LINE Notify API, to send visitor information and response content as push notifications to the user's smartphone. The input data consists of the response content and visitor information, and the output data is the push notification message sent to the user's device.

[1060] Step 9:

[1061] The user generates a follow-up message as needed and sends it to the server.

[1062] Specific operation: The user receives a push notification on their smartphone, enters a reply message as needed, and sends it to the server. The input data is the user's reply message, and the output data is the reply data sent to the server.

[1063] Step 10:

[1064] The server receives a response message from the user, converts it into an appropriate format using generative artificial intelligence, and transmits it to the visitor via the terminal.

[1065] Specific operation: The server receives a response message from the user, converts it into a voice message using generative artificial intelligence, and plays it back to the visitor through the terminal's speaker. The input data is the user's response message, and the output data is the voice message to the visitor.

[1066] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1067] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[1068] Initial processing when the intercom rings

[1069] Terminal:

[1070] When a visitor arrives at the intercom, a sensor detects this and sends an event signal to the server. This allows the server to recognize the visitor's presence.

[1071] server:

[1072] Upon receiving the event signal, the server activates the generative artificial intelligence and initiates an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[1073] Dialogue with visitors

[1074] Terminal:

[1075] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[1076] server:

[1077] The received audio data is converted into text data using speech recognition technology. This text data accurately reflects the content of the visitor's speech.

[1078] server:

[1079] The converted text data is analyzed using natural language processing to classify the visitor's purpose into multiple categories. Examples of these categories include: unattended delivery, solicitation, remote meeting, and childcare. Based on the analysis results, an appropriate automated response message is generated by a generative artificial intelligence system.

[1080] Automated responses and user notifications

[1081] server:

[1082] The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. This ensures that the visitor receives an appropriate response.

[1083] server:

[1084] The generated response and information regarding the visitor's purpose are sent to the user's device using a notification system.

[1085] User emotion recognition and adjustment

[1086] server:

[1087] An emotion engine built into the user's device recognizes emotions from the user's text or voice input. Based on this emotional state, notification content can be adjusted. For example, if the user is stressed, notifications will be delivered using softer language.

[1088] server:

[1089] Furthermore, the emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[1090] User response

[1091] User:

[1092] Users who receive notifications on devices such as smartphones can generate additional instructions or response messages via the server as needed. For example, if a friend comes to visit, they can send a custom message such as, "I'll be right back, please wait a moment."

[1093] Specific example

[1094] Example 1: Leaving delivery unattended when the recipient is absent.

[1095] 1. Terminal: The intercom rings.

[1096] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[1097] 3. Terminal: The visitor responds, "Delivery. I've come to deliver your package."

[1098] 4. Server: Receives voice data and converts it to text. Analyzes it using natural language processing to determine whether to leave the package unattended when the recipient is absent.

[1099] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[1100] 6. Terminal: Provides instructions to visitors.

[1101] 7. Server: Notifies the user that "Your delivery has arrived. The package has been left in front of your door."

[1102] 8. Server: If the emotion engine detects the user's stress level, it adjusts the notification content to a softer tone.

[1103] Example 2: Solicitation / Sales Strategies

[1104] 1. Terminal: The intercom rings.

[1105] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[1106] 3. Terminal: The visitor responds, "This is to inform you about a new service."

[1107] 4. Server: Receives audio data, converts it to text, analyzes it using natural language processing, and determines whether it is a solicitation / sales call.

[1108] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[1109] 6. Terminal: Provides instructions to visitors.

[1110] 7. Server: Notifies the user that "a solicitation was received and declined."

[1111] 8. Server: The emotion engine dynamically adjusts the tone and content of notifications according to the user's situation.

[1112] In this way, it is possible to efficiently handle interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[1113] The following describes the processing flow.

[1114] Step 1:

[1115] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[1116] Step 2:

[1117] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[1118] Step 3:

[1119] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[1120] Step 4:

[1121] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[1122] Step 5:

[1123] Terminal: Sends recorded audio data to the server via the internet.

[1124] Step 6:

[1125] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[1126] Step 7:

[1127] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[1128] Step 8:

[1129] Server: Based on the analysis results, the AI ​​generates an appropriate automated response message. This message includes the most suitable response for the request.

[1130] Step 9:

[1131] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[1132] Step 10:

[1133] Terminal: The response message is properly conveyed to the visitor.

[1134] Step 11:

[1135] Server: Sends the generated response and information about the visitor's purpose to the user's terminal using a notification system.

[1136] Step 12:

[1137] User: The user receives notifications on a device such as a smartphone, and the emotion engine analyzes the user's emotions.

[1138] Step 13:

[1139] Server: The emotion engine recognizes the user's emotional state and adjusts the notification content as needed. For example, if the user is feeling stressed, a notification will be sent in softer language.

[1140] Step 14:

[1141] Server: The emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[1142] Step 15:

[1143] User: Check the notification, generate a reply message from your smartphone if necessary, and transmit it to the visitor via the server. For example, you can send a custom message such as, "I'll be right back, please wait a moment."

[1144] These steps enable the system to efficiently automate visitor interactions while considering the user's emotional state, thereby improving user safety and convenience.

[1145] (Example 2)

[1146] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1147] Traditional intercom systems require manual responses to visitors, making it difficult to provide appropriate care when the user is absent or busy. Furthermore, they lack the ability to respond flexibly to emotions and situations, hindering user convenience and safety. Additionally, they often fail to accurately identify the visitor's purpose and situation, leading to delays in responses.

[1148] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1149] In this invention, the server includes a sensor means for detecting visitors, a computing means for activating generative artificial intelligence to conduct initial dialogue, and a speech recognition means for converting speech to text. This enables automatic detection of visitors and appropriate initial dialogue. It also includes a natural language processing means for analyzing text and classifying the visitor's request, and an emotion engine for recognizing the user's emotions and adjusting notification content and responses. This enables flexible responses that take into account the user's emotional state, improving user safety and convenience.

[1150] A "visitor" is a person who visits via the intercom system.

[1151] "Sensor means" refers to devices or means used to detect the presence of visitors, and includes motion detection and infrared sensors.

[1152] "Computer means" refers to a system or device that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[1153] "Generative artificial intelligence" refers to artificial intelligence technology that generates voice dialogues and response messages within a system.

[1154] "Voice recognition means" refers to technologies and devices for converting a visitor's voice into text data.

[1155] "Natural language processing means" refers to technologies and devices that analyze text data and classify the purpose of a visitor's visit.

[1156] A "response means" refers to a device or technology that generates an automated response message according to the classified request and transmits it to the visitor.

[1157] "Notification means" refers to technologies and devices used to notify the user's terminal of the content of the response or the purpose of the visitor's visit.

[1158] An "emotion engine" is a technology or system that recognizes a user's emotions and adjusts notification content and responses accordingly.

[1159] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, computer means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[1160] Initial processing when the intercom rings

[1161] When the terminal detects a visitor, the sensor detects this and sends an event signal to the server. Upon receiving the event signal, the server activates generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who are you?", "What can I help you with?", and "Who in your family is this for?". The generative artificial intelligence uses speech synthesis technology to output this as speech.

[1162] Dialogue with visitors

[1163] The device's microphone captures the visitor's responses and records them in real time. The recorded audio data is sent to a server. The server converts the received audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). After conversion, the text data is analyzed by a natural language processing unit (e.g., Hugging Face's NLP model) to classify the visitor's request into multiple categories.

[1164] Automated responses and user notifications

[1165] The server generates an appropriate automated response message based on the analysis results. The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. In addition, information regarding the response content and the visitor's purpose is sent to the user's terminal using notification methods (e.g., push notification to a smartphone or email).

[1166] User emotion recognition and adjustment

[1167] The server has an emotion engine that can recognize emotions from text and voice input from the user's device. The emotion engine utilizes the Emotion API and other tools to adjust notification content to a softer tone if the user is feeling stressed. The emotion engine can also dynamically adjust the tone and content of automated response messages based on the user's emotional state. If the user is relaxed, messages will be generated in a friendly tone.

[1168] User response

[1169] When a user receives a notification on a device such as a smartphone, they can generate additional instructions or reply messages via the server as needed. For example, if a friend visits, they can send a custom message such as, "I'll be right back, please wait a moment."

[1170] Specific example

[1171] Example 1: Leaving delivery unattended when the recipient is absent.

[1172] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[1173] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[1174] 3. Terminal: The intercom's microphone captures the visitor's speech, such as "Delivery. I've come to deliver your package," and sends it to the server.

[1175] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[1176] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if the package was left unattended when the recipient was absent.

[1177] 6. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[1178] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[1179] 8. Server: Notifies the user, "The delivery has arrived. The package has been left in front of your door."

[1180] 9. Server: The server uses the Emotion API to sense the user's stress level and adjusts the notification content to a softer tone.

[1181] Example of a prompt:

[1182] "If your doorbell rings and a visitor asks you to leave a package, how would you respond?"

[1183] Example 2: Solicitation / Sales Strategies

[1184] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[1185] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[1186] 3. Terminal: The intercom's microphone captures the visitor's statement, "We have an announcement about a new service," and sends it to the server.

[1187] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[1188] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if it is a solicitation / sales pitch.

[1189] 6. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[1190] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[1191] 8. Server: Notifies the user that "a solicitation was received and declined."

[1192] 9. Server: Use the Emotion API to check the user's status and dynamically adjust the tone and content of notifications.

[1193] In this way, the system efficiently handles interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[1194] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1195] Step 1:

[1196] Terminal: The intercom's sensor detects a visitor. As input, the sensor obtains information about the visitor's presence. When the sensor is triggered, an event signal is generated and sent to the server. As output, the server recognizes the visitor's presence.

[1197] Step 2:

[1198] Server: The server receives an event signal and activates the generative artificial intelligence. It receives the event signal as input. The generative artificial intelligence generates initial dialogue questions. It generates questions such as "Who is this?", "What can I help you with?", and "Who in your family is this addressed to?" as text and converts them into speech using speech synthesis technology. It generates the voiced questions as output and sends them to the terminal.

[1199] Step 3:

[1200] Terminal: The intercom's microphone captures and records the visitor's response. It receives the visitor's voice response as input. The recorded data is sent to the server. It also sends the voice data to the server as output.

[1201] Step 4:

[1202] Server: The server converts the received audio data into text data using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). It receives audio data as input. The speech recognition engine analyzes the audio and generates text data. It generates text data as output.

[1203] Step 5:

[1204] Server: The converted text data is analyzed using natural language processing (e.g., Hugging Face's NLP model). It receives text data as input. The natural language processing analyzes the text content and classifies the visitor's request into multiple categories. It generates classified category information as output.

[1205] Step 6:

[1206] Server: Generates an appropriate automated response message based on the analysis results. It receives classified category information as input. The generative artificial intelligence generates a response message such as, "We are currently out, please leave your package at the front door." It sends the generated response message to the terminal as output.

[1207] Step 7:

[1208] Terminal: The terminal transmits the response message received from the server to the visitor through the intercom speaker. Input: Receives the response message. Plays the voice message using the intercom speaker. Output: Transmits the response message to the visitor.

[1209] Step 8:

[1210] Server: The generated response and information regarding the visitor's request are sent to the user's device using a notification method. The server receives the response and request information as input. It then notifies the user's device using a notification method (e.g., push notification or email). The notification is then delivered to the user as output.

[1211] Step 9:

[1212] Server: The server contains an emotion engine that recognizes the user's emotions from text and voice input from the user's device. It receives user input data as input. It analyzes the user's emotional state using the Emotion API and other tools. It generates user emotion information as output.

[1213] Step 10:

[1214] Server: The emotion engine adjusts notification content based on the user's emotional state. It receives user emotion information as input. It adjusts the notification content and tone according to the emotional state to generate an appropriate notification message. As output, it sends the adjusted notification message to the user's device.

[1215] Step 11:

[1216] User: The user receives a notification on a device such as a smartphone and generates additional instructions or reply messages via the server as needed. Input: Receives the user notification. Creates a custom message (e.g., "I'll be right back, please wait a moment") as needed and sends it to the server. Output: The reply message is transmitted to the visitor.

[1217] (Application Example 2)

[1218] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1219] Traditional intercom systems often required manual interaction with visitors, making quick and appropriate responses difficult, especially when users were away from home, working remotely, or with children home alone. Furthermore, the lack of appropriate notifications and responses tailored to the user's emotional state often led to inconvenience and stress. As a result, many users felt a lack of both safety and convenience.

[1220] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to perform an initial dialogue, a means for the generative artificial intelligence to convert speech into text, a means for analyzing the text and classifying the visitor's request, a means for generating an automatic response message according to the classified request and transmitting it to the visitor, a means for notifying the request and response content to the user's communication terminal, and a means for recognizing the user's emotions and adjusting the tone and content of the notification and automatic response message. This enables quick and appropriate responses even when the user is absent or busy, and allows for flexible notifications and responses according to the user's emotional state.

[1221] "Sensor means" refers to a device or system used to detect visitors.

[1222] A "server means" is a device or system that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[1223] "Generative artificial intelligence" refers to artificial intelligence software that converts speech into text, analyzes that text, and generates appropriate response messages.

[1224] "Voice recognition means" refers to a device or system that converts a visitor's voice into text.

[1225] A "natural language processing device" is a device or system that analyzes text to classify the visitor's request.

[1226] "Response means" refers to a device or system that generates an automated response message according to the classified requirements and transmits it to the visitor.

[1227] "Notification means" refers to a device or system that notifies the user's communication terminal of the aforementioned request and response content.

[1228] "Emotion recognition means" refers to a device or system that recognizes the user's emotions and adjusts the tone and content of the notification and automated response messages.

[1229] A "communication terminal" is an electronic device, including smartphones and tablets, that a user uses to receive notifications.

[1230] The system for implementing this invention automatically performs a series of operations from visitor detection to response and notification, and further recognizes the user's emotions to adjust notifications and responses accordingly. The specific form of this system is described below.

[1231] Hardware configuration

[1232] This system uses the following hardware:

[1233] Sensory devices (e.g., intercoms, security cameras, door sensors)

[1234] Server configuration (e.g., cloud server)

[1235] User's communication device (e.g., smartphone, tablet)

[1236] Software to use

[1237] This system is equipped with the following software:

[1238] Generative artificial intelligence (e.g. GPT-3)

[1239] Speech recognition methods (e.g., Google Speech Recognition API)

[1240] Natural language processing tools (e.g., TextBlob library)

[1241] Emotion recognition methods (e.g., TextBlob's sentiment analysis function)

[1242] Data processing and data calculation

[1243] The server performs the following data processing and calculations.

[1244] 1. Visitor detection

[1245] Sensors detect the presence of visitors. Examples include intercoms and door sensors.

[1246] 2. Initiating the initial dialogue

[1247] The server receives a notification from a sensor and activates a generative artificial intelligence (AI) to initiate an initial conversation with the visitor, such as "Who are you?". This AI then uses speech recognition to convert the visitor's voice into text.

[1248] 3. Classification of Visitor's Purpose

[1249] Text data converted by generative artificial intelligence is analyzed using natural language processing techniques to appropriately classify the visitor's purpose.

[1250] 4. Generation and transmission of response messages

[1251] Based on the categorized request, an automated response message is generated and communicated to the visitor via the response system.

[1252] 5. Notification to the user's terminal

[1253] Simultaneously, the server sends information about the response and the visitor's purpose to the user's communication terminal using a notification mechanism.

[1254] 6. User emotion recognition and notification adjustment

[1255] Furthermore, using emotion recognition mechanisms, the system recognizes the user's emotions from their text or voice input and dynamically adjusts the tone of notification content and response messages.

[1256] Specific example

[1257] The following are specific examples.

[1258] Example 1: Delivery company's response

[1259] 1. The sensor detects the delivery person.

[1260] 2. The server activates a generative artificial intelligence and responds, "Who is this?"

[1261] 3. The delivery person responds, "I've come to deliver your package."

[1262] 4. The speech recognition means converts the speech into text, and the natural language processing means analyzes the text.

[1263] 5. The server generates a response message saying, "We are currently out, please leave it at the front door," and transmits it to the delivery person via the response system.

[1264] 6. The user's communication device is notified with the message, "Your delivery has arrived. The package has been left in front of your door."

[1265] Example 2: Detection of a suspicious person

[1266] 1. The sensor detects a suspicious person.

[1267] 2. The server activates a generative artificial intelligence and responds, "Who is it?", but the suspicious person remains silent.

[1268] 3. The server identifies the person as suspicious, generates a warning message, and transmits it via the response system.

[1269] 4. Simultaneously, a warning notification is sent to the user's communication terminal.

[1270] Example of a prompt

[1271] "Person says: I'm here to deliver a package. User is stressed. Generate a soft response."

[1272] In this way, it becomes possible to respond quickly and appropriately even when the user is away or busy, and to provide flexible notifications and responses that are tailored to their emotional state.

[1273] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1274] Step 1:

[1275] A sensor detects the presence of a visitor. The input is the sensor (e.g., intercom, security camera, door sensor), and the output is a notification signal of the detection event. This notification signal is sent to the server.

[1276] Step 2:

[1277] The server receives a notification from a sensor. The input is a notification signal, and the server activates a generative artificial intelligence to initiate an initial conversation. Specifically, it generates questions such as "Who are you?" for the visitor and outputs them as voice through a response system.

[1278] Step 3:

[1279] The visitor's response is captured by the intercom's microphone and recorded in real time as audio data. The input is the visitor's voice, and the output is the recorded audio data. The terminal sends this audio data to the server.

[1280] Step 4:

[1281] The server converts the received audio data into text data using speech recognition tools (e.g., Google Speech Recognition API). The input is audio data, and the output is text data. Specifically, the speech recognition engine analyzes the audio waveform data and converts it into the corresponding text.

[1282] Step 5:

[1283] The server analyzes the converted text data using natural language processing tools (e.g., the TextBlob library). The input is text data, and the output is the category of requests resulting from the analysis. In this process, the content of the text is analyzed, and its intent and requirements are classified into categories such as receiving packages when absent, refusing solicitations, responding during remote meetings, and how to handle situations when children are home alone.

[1284] Step 6:

[1285] The server generates an automated response message based on the analysis results. The input is the category of the request, and the output is the automated response message. Specifically, a generative artificial intelligence generates an appropriate message and conveys it to the visitor. For example, if the recipient is absent and the package is to be left at the door, the server will generate a message such as, "We are currently absent, please leave your package at the front door."

[1286] Step 7:

[1287] The server transmits a generated response message to the visitor via the response mechanism. The input is an automated response message, and the output is an audio message to the visitor. Specifically, it is transmitted to the visitor through the intercom speaker.

[1288] Step 8:

[1289] The server sends the generated response and information about the visitor's purpose to the user's communication terminal using a notification mechanism. The input is the response and purpose information, and the output is a notification message to the user's terminal. Specifically, it is displayed as a notification on the user's smartphone.

[1290] Step 9:

[1291] The server uses emotion recognition to recognize the user's emotions and adjusts the tone and content of notifications and automated response messages accordingly. The input is the user's emotion data (e.g., text or voice input), and the output is the adjusted notification and response message. For example, if the user is feeling stressed, the notification content might be adjusted to a softer tone, such as "Your package has been left at your front door."

[1292] Step 10:

[1293] The user generates additional instructions or follow-up messages as needed, which are then transmitted to the visitor via the server. The input is the user's follow-up message, and the output is a voice message to the visitor. For example, the user might send a custom message such as, "I'll be right back, please wait a moment."

[1294] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1295] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1296] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1297] [Fourth Embodiment]

[1298] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1299] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1300] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1301] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1302] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1303] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1304] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1305] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1306] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1307] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1308] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1309] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1310] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1311] The system for carrying out this invention automatically processes a series of operations from visitor detection to response and notification. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, and notification means.

[1312] Initial processing when the intercom rings

[1313] Terminal:

[1314] When a visitor arrives at the intercom, a sensor detects them and sends an event signal to the server. This allows the server to recognize that a visitor has arrived.

[1315] server:

[1316] Upon receiving the event signal, the server activates the generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[1317] Dialogue with visitors

[1318] Terminal:

[1319] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[1320] server:

[1321] The server converts the received audio data into text data using speech recognition. The converted text data is then analyzed by natural language processing to classify the visitor's request.

[1322] Classification of inquiries and automated responses

[1323] server:

[1324] Natural language processing tools analyze text data and classify visitors' requests into several categories. These categories include, for example, unattended delivery, solicitation or sales calls, remote meetings, and children home alone. An automated response message tailored to each request is then generated by a generative artificial intelligence system.

[1325] Terminal:

[1326] The generated response message is transmitted to the visitor through the device's speaker, ensuring they receive an appropriate response.

[1327] Notification to the user

[1328] server:

[1329] The generated response and information regarding the visitor's purpose are sent as push notifications to the user's device using a notification system. For example, if the user has a smartphone, a specific notification such as "We have confirmed that Sagawa Express has arrived and left the package at your front door" will be sent.

[1330] User:

[1331] Users can review received notifications, generate a follow-up message if necessary, and transmit it to visitors via the server. This follow-up allows users to send custom messages, such as "I'll be right back, please wait a moment," when a friend visits.

[1332] Specific example

[1333] Here are some specific examples.

[1334] Example 1: Leaving delivery unattended when the recipient is absent.

[1335] 1. Terminal: The intercom rings.

[1336] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[1337] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver a package."

[1338] 4. Server: Receives voice data, converts it to text, and analyzes it using natural language processing. Determines that it is a delivery attempt when the recipient is absent.

[1339] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[1340] 6. Terminal: Provides instructions to visitors.

[1341] 7. Server: Notifies the user that "A delivery service has arrived. The package has been left in front of your door."

[1342] Example 2: Dealing with persistent solicitation / sales tactics

[1343] 1. Terminal: The intercom rings.

[1344] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[1345] 3. Terminal: The visitor responds, "I've come to introduce a new service."

[1346] 4. Server: Receives audio data and converts it to text. Analysis using natural language processing techniques indicates that it is a solicitation / sales pitch.

[1347] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[1348] 6. Terminal: Provides instructions to visitors.

[1349] This system can streamline intercom communication in users' daily lives, improving security and convenience.

[1350] The following describes the processing flow.

[1351] Step 1:

[1352] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[1353] Step 2:

[1354] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[1355] Step 3:

[1356] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[1357] Step 4:

[1358] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[1359] Step 5:

[1360] Terminal: Sends recorded audio data to the server via the internet.

[1361] Step 6:

[1362] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[1363] Step 7:

[1364] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[1365] Step 8:

[1366] Server: Based on the classification results, an appropriate automated response message is generated by the AI ​​generation system. This message includes the optimal response method for the given request.

[1367] Step 9:

[1368] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[1369] Step 10:

[1370] Terminal: The response message is properly conveyed to the visitor, and the response ends.

[1371] Step 11:

[1372] Server: Notifies the user's terminal of the generated response and information regarding the visitor's purpose.

[1373] Step 12:

[1374] Users can receive notifications on devices such as smartphones, and, if necessary, generate additional instructions or response messages via the server and communicate them to visitors.

[1375] These steps automate the intercom system, allowing for efficient handling of visitors without direct user interaction.

[1376] (Example 1)

[1377] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1378] Traditional intercom systems require users to manually handle a series of processes, from detecting visitors to answering, providing appropriate responses, and notifying the user, which places a significant burden on the user. This inconvenience is particularly noticeable when users are away for extended periods or working remotely. Furthermore, dealing with inappropriate solicitations and sales calls presents challenges, highlighting the need for automated systems to address these issues.

[1379] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1380] In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to conduct an initial dialogue, a speech recognition means in which the generative artificial intelligence converts speech into text, a natural language processing means for analyzing the text and classifying the visitor's request, a response means for generating an automated response message according to the classified request and communicating it to the visitor, and a notification means for notifying the user's communication device of the request and response content. As a result, a series of processes from visitor detection to response and notification are automated, significantly reducing the burden on the user, improving convenience, and ensuring security.

[1381] A "sensor device" is a device that detects visitors, generates an event signal, and notifies the server.

[1382] "Generative artificial intelligence" is a program that generates text questions and answers to automate interactions with visitors.

[1383] "Speech recognition means" refers to a device or program that converts received speech data into text data.

[1384] A "natural language processing tool" is a program that analyzes text data, understands its meaning, and classifies it into appropriate categories.

[1385] A "response means" refers to a device or program for transmitting a generated response message to an visitor.

[1386] A "notification means" is a device or program that notifies the user's communication device of the generated response content or information regarding the visitor's purpose.

[1387] A "server system" is a computer system that activates generative artificial intelligence and controls and manages various systems in a centralized manner.

[1388] "Communication devices" refer to terminals that allow users to receive notifications, and include smartphones and tablets.

[1389] The system for implementing this invention automatically processes a series of actions from visitor detection to response and notification. The hardware includes sensors, terminals, servers, and communication equipment, while the software includes generative artificial intelligence, a speech recognition engine, a natural language processing engine, and a notification system.

[1390] First, the system uses sensors to detect visitors. For example, infrared sensors or motion sensors may be used. When a visitor is detected, the sensor generates an event signal, which is transmitted to the server via the network.

[1391] When the server receives this event signal, it activates a generative artificial intelligence (e.g., the GPT-3 model). The generative AI then generates prompt sentences to start the initial conversation. For example, it might generate questions such as "Who is this?", "What can I help you with?", or "Who in your family is this addressed to?". The generated prompt sentences are then converted into speech data. A text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) performs this conversion.

[1392] The generated audio data is played back through the terminal's speaker, conveying the question to the visitor. Next, the intercom's microphone records the visitor's response in real time and sends the audio data to the server. The server uses speech recognition (e.g., IBM Watson Speech to Text API) to convert the received audio data into text data.

[1393] The converted text data is analyzed using natural language processing tools (e.g., the spaCy library), and the visitor's request is classified into several categories. For example, possible categories include receiving packages when absent, refusing solicitations, responding during remote meetings, and handling children left alone at home.

[1394] Based on the classification results, the generative artificial intelligence generates an appropriate response message. For example, a message such as, "I am currently out, please leave your package in front of the door." This response message is sent to the terminal and conveyed to the visitor through the intercom speaker. The generated response content and the visitor's request are recorded on the server and also pushed to the user's communication device using a notification method (e.g., Firebase Cloud Messaging).

[1395] Users can check notifications on communication devices such as smartphones and generate a follow-up message if necessary. By sending this follow-up message to the server, it becomes possible to provide visitors with appropriate instructions again.

[1396] Specific example

[1397] Example 1: Leaving delivery unattended when the recipient is absent.

[1398] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[1399] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[1400] 3. Terminal: The visitor responds, "This is Sagawa Express. I've come to deliver your package," and sends the voice data to the server.

[1401] 4. Server: Uses speech recognition to convert speech data into text data, which is then analyzed using natural language processing. This is classified as "delivery left unattended when the recipient is absent."

[1402] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[1403] 6. Terminal: Receives a response message from the visitor via the speaker.

[1404] 7. Server: Notifies the user's communication device with the message, "A delivery service has arrived. The package has been left in front of your door."

[1405] Example 2: Dealing with persistent solicitation / sales tactics

[1406] 1. Server: The generative artificial intelligence generates questions such as, "Who is this? What is your purpose? Who in your family is this addressed to?" and uses a text-to-speech engine to create audio data.

[1407] 2. Device: Plays audio data through a speaker to convey questions to visitors.

[1408] 3. Terminal: The visitor responds, "I've come to introduce you to a new service," and sends the audio data to the server.

[1409] 4. Server: Converts speech data into text data using speech recognition technology and analyzes it using natural language processing technology. Classified as solicitation / sales.

[1410] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[1411] 6. Terminal: Receives a response message from the visitor via the speaker.

[1412] This system allows users to significantly automate their intercom responses, improving the convenience and safety of their daily lives.

[1413] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1414] Step 1:

[1415] The terminal detects visitors using the intercom's sensor. The sensor consists of infrared sensors and motion sensors, and generates an event signal when it detects a visitor's movement. The input is the presence of a visitor, and the output is the transmission of the event signal to the server.

[1416] Step 2:

[1417] The server receives an event signal from a sensor. Upon receiving the event signal, the server activates a generative artificial intelligence (e.g., the GPT-3 model) to generate a prompt for initial interaction. The input is the event signal, and the output is the prompt. Specifically, the generative artificial intelligence generates a question such as "Who are you?"

[1418] Step 3:

[1419] The server uses a text-to-speech (TTS) engine (e.g., Google Text-to-Speech API) to convert the generated prompt text into speech data. The input is the prompt text, and the output is speech data. The TTS engine outputs the prompt text as synthesized speech.

[1420] Step 4:

[1421] The terminal receives audio data transmitted from the server and plays it back through the intercom speaker. The input is audio data, and the output is the audio message conveyed to the visitor. The speaker plays the audio data and conveys the question to the visitor.

[1422] Step 5:

[1423] The terminal uses the intercom's microphone to record the visitor's response in real time. The recorded audio data is immediately sent to the server. The input is the visitor's voice, and the output is the transmission of audio data to the server.

[1424] Step 6:

[1425] The server converts the received audio data into text data using speech recognition software (e.g., IBM Watson Speech to Text API). The input is audio data, and the output is text data. The speech recognition software analyzes the audio and outputs it as text data.

[1426] Step 7:

[1427] The server sends text data to a natural language processing system (e.g., the spaCy library), which analyzes the content and classifies the visitor's request. The input is text data, and the output is the classification result. The natural language processing system performs analysis and classification, assigning the request to a specific category (e.g., receiving a package when absent, refusing a solicitation, etc.).

[1428] Step 8:

[1429] The server uses generative artificial intelligence to generate an appropriate response message based on the classification result. The input is the classification result, and the output is the response message. The generative artificial intelligence then generates a prompt sentence, which is used as the response message.

[1430] Step 9:

[1431] The terminal receives a response message (audio data) from the server and transmits it to the visitor through the intercom speaker. The input is the response message, and the output is the audio transmitted to the visitor. The speaker plays the response message and responds to the visitor.

[1432] Step 10:

[1433] The server records the response and visitor's request and sends a push notification to the user's device using a notification method (e.g., Firebase Cloud Messaging). The input is the response and visitor's request, and the output is a push notification to the user's device. The notification method generates the message and sends it to the user's device.

[1434] Step 11:

[1435] The user can review the notification and, if necessary, generate a reply message and send it to the server. The input is the user's reply message, and the output is the message sent to the server. The user inputs and sends the message using a smartphone or similar device.

[1436] Step 12:

[1437] The server processes the received reply message and transmits it to the visitor through the terminal. The input is the reply message, and the output is the audio message to be delivered to the visitor. The server converts the reply message into audio data and plays it through the terminal's speaker.

[1438] The above outlines the specific processing steps of this system. This automates the entire process from visitor detection to response and notification, significantly reducing the burden on the user.

[1439] (Application Example 1)

[1440] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1441] In recent years, many homes and businesses have frequently had to deal with visitors, and prompt and appropriate responses are especially required when people are away or busy. However, current intercom systems require manual responses, which is inefficient and poses security risks. Furthermore, it is difficult to respond quickly to unexpected visitors such as solicitors or salespeople, resulting in wasted time and effort for users. Moreover, if these responses are not handled properly, important notifications may be delayed, which can negatively impact users' quality of life.

[1442] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1443] In this invention, the server includes detection means for detecting visitors, computation means for activating generative artificial intelligence to conduct initial dialogue, speech recognition means for converting speech to text, natural language processing means for analyzing text and classifying the visitor's request, response means for generating and responding with an automated response message, speech synthesis means for generating a voice response, and push notification means for sending visitor information to the user's terminal. This makes it possible to automate visitor interactions efficiently and securely. Furthermore, it enables appropriate responses in various situations such as handling deliveries when the user is absent and refusing solicitations, reducing the burden on the user and improving security.

[1444] A "visitor" is a person who visits a specific location, such as a home or office, and attempts to make contact using an intercom.

[1445] A "detection means" is a device or sensor used to detect the presence of a visitor, and it plays the first role in recognizing a visitor.

[1446] "Generative artificial intelligence" refers to artificial intelligence technologies that automatically perform intelligent tasks such as conversation, text generation, and speech recognition.

[1447] "Computation means" refers to a part of a server or computer system that receives and processes data from sensors and other devices.

[1448] "Voice recognition means" refers to a technology or device that converts voice input into text data, and is a means of recording the content of conversations with visitors as text.

[1449] "Natural language processing" refers to technologies that classify and understand specific intentions and requirements through the analysis of text data.

[1450] A "response means" is a technology or device for generating an appropriate response message based on classified requirements and communicating it to visitors.

[1451] A "speech synthesis engine" is a technology that converts text data into speech data and outputs it as speech.

[1452] "Push notification methods" refer to technologies for sending notifications to a user's device in real time.

[1453] This invention is a system that automatically processes a series of actions from visitor detection to response and notification. This system mainly consists of the following components: a detection means for detecting visitors, a computation means for activating generative artificial intelligence and conducting an initial dialogue, a speech recognition means for converting speech to text, a natural language processing means for analyzing text data and classifying the visitor's request, a response means for generating a response message and communicating it to the visitor, a push notification means for sending visitor information to the user's terminal, and a speech synthesis engine.

[1454] When the server receives a signal from the sensor, it activates generative artificial intelligence and begins an initial conversation with the visitor. For example, it might ask questions such as, "Who is this? What can I help you with? Who in your family is this for?" The initial conversation is generated by a speech synthesis engine and output through the speaker. The visitor's responses are captured by the intercom microphone and converted into text through speech recognition.

[1455] Next, the server analyzes the converted text using natural language processing to classify the visitor's request. Based on the classified request, a generative artificial intelligence generates an appropriate response message and transmits it to the visitor via a response system.

[1456] For example, if a visitor responds, "It's a delivery service. I've come to deliver your package," the server generates a response message saying, "You are currently out, please leave your package in front of the door," and outputs it using a speech synthesis engine.

[1457] Furthermore, the server uses push notifications to inform the user's device of visitor information and response content. This notification is delivered in real time using APIs such as LINE Notify API. Users receive notifications through their devices and can generate custom response messages as needed, then transmit them to the visitor again via the server.

[1458] The following scenarios are possible as specific examples.

[1459] When a visitor presses the intercom, the following notification is sent to the smartphone: "Who is it? What can I do for you? Who in the family is this for?". If the visitor replies, "It's a delivery. I've come to deliver a package," the system responds, "You are currently out, please leave the package in front of the door," and simultaneously sends a notification to the user's smartphone saying, "A delivery person has arrived. They have left the package in front of the door."

[1460] In this way, this invention automates visitor reception, significantly reducing the burden on users and improving security.

[1461] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1462] Step 1:

[1463] The sensor detects the visitor.

[1464] Specific operation: The detection device (sensor) detects when a visitor presses the intercom and sends an event signal to the server. The input data is information about the visitor's presence, and the output data is the event signal sent to the server.

[1465] Step 2:

[1466] The server receives an event signal and activates the generative artificial intelligence.

[1467] Specific operation: The server receives an event signal from the sensor and activates a generative artificial intelligence (AI). The generative AI prepares initial questions and uses a speech synthesis engine to output a voice message such as, "Who is this? What can I help you with? Who in your family is this for?" The input data is the event signal, and the output data is a voice message to the visitor.

[1468] Step 3:

[1469] The device captures the visitor's response and sends it to the server.

[1470] Specific operation: The intercom's microphone captures the visitor's response and sends it to the server as audio data. The input data is the visitor's voice, and the output data is the audio data sent to the server.

[1471] Step 4:

[1472] The server receives the audio data and converts it into text using speech recognition technology.

[1473] Specific operation: The server converts audio data into text data using speech recognition technology. The input data is audio data, and the output data is the converted text data.

[1474] Step 5:

[1475] The server analyzes text data using natural language processing to classify the visitor's purpose.

[1476] Specific operation: The server analyzes text data using natural language processing and classifies it into categories such as "delivery" and "sales." Input data is text data, and output data is classified category information.

[1477] Step 6:

[1478] The server uses generative artificial intelligence to generate response messages based on classified requests, and a speech synthesis engine creates voice messages.

[1479] Specific operation: The server uses generative artificial intelligence to generate appropriate response messages for classified requests and outputs them as voice messages using a speech synthesis engine. For example, if a visitor says "I have a delivery," the server will generate a message such as "I am currently out, please leave your package in front of the door." The input data is classified request information, and the output data is the response voice message.

[1480] Step 7:

[1481] The terminal transmits the generated response message to the visitor.

[1482] Specific operation: The generated response message is played to the visitor through the device's speaker. The input data is the response voice message, and the output data is the voice output to the visitor.

[1483] Step 8:

[1484] The server creates and sends a push notification to the user's device containing visitor information and response details.

[1485] Specific operation: The server uses push notification methods, such as the LINE Notify API, to send visitor information and response content as push notifications to the user's smartphone. The input data consists of the response content and visitor information, and the output data is the push notification message sent to the user's device.

[1486] Step 9:

[1487] The user generates a follow-up message as needed and sends it to the server.

[1488] Specific operation: The user receives a push notification on their smartphone, enters a reply message as needed, and sends it to the server. The input data is the user's reply message, and the output data is the reply data sent to the server.

[1489] Step 10:

[1490] The server receives a response message from the user, converts it into an appropriate format using generative artificial intelligence, and transmits it to the visitor via the terminal.

[1491] Specific operation: The server receives a response message from the user, converts it into a voice message using generative artificial intelligence, and plays it back to the visitor through the terminal's speaker. The input data is the user's response message, and the output data is the voice message to the visitor.

[1492] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1493] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, server means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[1494] Initial processing when the intercom rings

[1495] Terminal:

[1496] When a visitor arrives at the intercom, a sensor detects this and sends an event signal to the server. This allows the server to recognize the visitor's presence.

[1497] server:

[1498] Upon receiving the event signal, the server activates the generative artificial intelligence and initiates an initial conversation. This conversation includes questions such as "Who is this?", "What is your purpose?", and "Who in your family is this addressed to?". These questions are output as speech by the generative artificial intelligence.

[1499] Dialogue with visitors

[1500] Terminal:

[1501] The intercom's microphone captures the visitor's response and records it in real time. The recorded audio data is then sent to a server.

[1502] server:

[1503] The received audio data is converted into text data using speech recognition technology. This text data accurately reflects the content of the visitor's speech.

[1504] server:

[1505] The converted text data is analyzed using natural language processing to classify the visitor's purpose into multiple categories. Examples of these categories include: unattended delivery, solicitation, remote meeting, and childcare. Based on the analysis results, an appropriate automated response message is generated by a generative artificial intelligence system.

[1506] Automated responses and user notifications

[1507] server:

[1508] The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. This ensures that the visitor receives an appropriate response.

[1509] server:

[1510] The generated response and information regarding the visitor's purpose are sent to the user's device using a notification system.

[1511] User emotion recognition and adjustment

[1512] server:

[1513] An emotion engine built into the user's device recognizes emotions from the user's text or voice input. Based on this emotional state, notification content can be adjusted. For example, if the user is stressed, notifications will be delivered using softer language.

[1514] server:

[1515] Furthermore, the emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[1516] User response

[1517] User:

[1518] Users who receive notifications on devices such as smartphones can generate additional instructions or response messages via the server as needed. For example, if a friend comes to visit, they can send a custom message such as, "I'll be right back, please wait a moment."

[1519] Specific example

[1520] Example 1: Leaving delivery unattended when the recipient is absent.

[1521] 1. Terminal: The intercom rings.

[1522] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[1523] 3. Terminal: The visitor responds, "Delivery. I've come to deliver your package."

[1524] 4. Server: Receives voice data and converts it to text. Analyzes it using natural language processing to determine whether to leave the package unattended when the recipient is absent.

[1525] 5. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[1526] 6. Terminal: Provides instructions to visitors.

[1527] 7. Server: Notifies the user that "Your delivery has arrived. The package has been left in front of your door."

[1528] 8. Server: If the emotion engine detects the user's stress level, it adjusts the notification content to a softer tone.

[1529] Example 2: Solicitation / Sales Strategies

[1530] 1. Terminal: The intercom rings.

[1531] 2. Server: The generative artificial intelligence outputs, "Who is this? What is your request? Which family member is this addressed to?"

[1532] 3. Terminal: The visitor responds, "This is to inform you about a new service."

[1533] 4. Server: Receives audio data, converts it to text, analyzes it using natural language processing, and determines whether it is a solicitation / sales call.

[1534] 5. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[1535] 6. Terminal: Provides instructions to visitors.

[1536] 7. Server: Notifies the user that "a solicitation was received and declined."

[1537] 8. Server: The emotion engine dynamically adjusts the tone and content of notifications according to the user's situation.

[1538] In this way, it is possible to efficiently handle interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[1539] The following describes the processing flow.

[1540] Step 1:

[1541] Terminal: The intercom's sensor detects a visitor and sends an event signal to the server.

[1542] Step 2:

[1543] Server: Upon receiving an event signal, it activates the generative artificial intelligence and prepares to begin the initial dialogue.

[1544] Step 3:

[1545] Server: The generative artificial intelligence uses a voice output mechanism to output questions such as "Who is this?", "What can I help you with?", and "Who in your family is this for?" through the intercom speaker.

[1546] Step 4:

[1547] Terminal: The intercom's microphone captures the visitor's voice response and records it in real time.

[1548] Step 5:

[1549] Terminal: Sends recorded audio data to the server via the internet.

[1550] Step 6:

[1551] Server: Converts received audio data into text data using speech recognition technology. This text data accurately reflects the visitor's spoken content.

[1552] Step 7:

[1553] Server: Passes the converted text data to a natural language processing system for text analysis. Classifies the requirements based on the analysis results.

[1554] Step 8:

[1555] Server: Based on the analysis results, the AI ​​generates an appropriate automated response message. This message includes the most suitable response for the request.

[1556] Step 9:

[1557] Server: Sends the generated response message to the terminal and transmits it to the visitor through the intercom speaker.

[1558] Step 10:

[1559] Terminal: The response message is properly conveyed to the visitor.

[1560] Step 11:

[1561] Server: Sends the generated response and information about the visitor's purpose to the user's terminal using a notification system.

[1562] Step 12:

[1563] User: The user receives notifications on a device such as a smartphone, and the emotion engine analyzes the user's emotions.

[1564] Step 13:

[1565] Server: The emotion engine recognizes the user's emotional state and adjusts the notification content as needed. For example, if the user is feeling stressed, a notification will be sent in softer language.

[1566] Step 14:

[1567] Server: The emotion engine can dynamically change the tone and content of automated response messages based on the user's emotional state. For example, if the user is relaxed, a message in a friendly tone will be generated.

[1568] Step 15:

[1569] User: Check the notification, generate a reply message from your smartphone if necessary, and transmit it to the visitor via the server. For example, you can send a custom message such as, "I'll be right back, please wait a moment."

[1570] These steps enable the system to efficiently automate visitor interactions while considering the user's emotional state, thereby improving user safety and convenience.

[1571] (Example 2)

[1572] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1573] Traditional intercom systems require manual responses to visitors, making it difficult to provide appropriate care when the user is absent or busy. Furthermore, they lack the ability to respond flexibly to emotions and situations, hindering user convenience and safety. Additionally, they often fail to accurately identify the visitor's purpose and situation, leading to delays in responses.

[1574] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1575] In this invention, the server includes a sensor means for detecting visitors, a computing means for activating generative artificial intelligence to conduct initial dialogue, and a speech recognition means for converting speech to text. This enables automatic detection of visitors and appropriate initial dialogue. It also includes a natural language processing means for analyzing text and classifying the visitor's request, and an emotion engine for recognizing the user's emotions and adjusting notification content and responses. This enables flexible responses that take into account the user's emotional state, improving user safety and convenience.

[1576] A "visitor" is a person who visits via the intercom system.

[1577] "Sensor means" refers to devices or means used to detect the presence of visitors, and includes motion detection and infrared sensors.

[1578] "Computer means" refers to a system or device that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[1579] "Generative artificial intelligence" refers to artificial intelligence technology that generates voice dialogues and response messages within a system.

[1580] "Voice recognition means" refers to technologies and devices for converting a visitor's voice into text data.

[1581] "Natural language processing means" refers to technologies and devices that analyze text data and classify the purpose of a visitor's visit.

[1582] A "response means" refers to a device or technology that generates an automated response message according to the classified request and transmits it to the visitor.

[1583] "Notification means" refers to technologies and devices used to notify the user's terminal of the content of the response or the purpose of the visitor's visit.

[1584] An "emotion engine" is a technology or system that recognizes a user's emotions and adjusts notification content and responses accordingly.

[1585] The system for carrying out this invention is characterized by automatically performing a series of operations from visitor detection to response and notification, and further by recognizing the user's emotions and adjusting notifications and responses accordingly. This system includes sensor means, computer means, generative artificial intelligence, speech recognition means, natural language processing means, response means, notification means, and emotion engine.

[1586] Initial processing when the intercom rings

[1587] When the terminal detects a visitor, the sensor detects this and sends an event signal to the server. Upon receiving the event signal, the server activates generative artificial intelligence and begins an initial conversation. This conversation includes questions such as "Who are you?", "What can I help you with?", and "Who in your family is this for?". The generative artificial intelligence uses speech synthesis technology to output this as speech.

[1588] Dialogue with visitors

[1589] The device's microphone captures the visitor's responses and records them in real time. The recorded audio data is sent to a server. The server converts the received audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). After conversion, the text data is analyzed by a natural language processing unit (e.g., Hugging Face's NLP model) to classify the visitor's request into multiple categories.

[1590] Automated responses and user notifications

[1591] The server generates an appropriate automated response message based on the analysis results. The generated response message is sent to the terminal and transmitted to the visitor through the intercom speaker. In addition, information regarding the response content and the visitor's purpose is sent to the user's terminal using notification methods (e.g., push notification to a smartphone or email).

[1592] User emotion recognition and adjustment

[1593] The server has an emotion engine that can recognize emotions from text and voice input from the user's device. The emotion engine utilizes the Emotion API and other tools to adjust notification content to a softer tone if the user is feeling stressed. The emotion engine can also dynamically adjust the tone and content of automated response messages based on the user's emotional state. If the user is relaxed, messages will be generated in a friendly tone.

[1594] User response

[1595] When a user receives a notification on a device such as a smartphone, they can generate additional instructions or reply messages via the server as needed. For example, if a friend visits, they can send a custom message such as, "I'll be right back, please wait a moment."

[1596] Specific example

[1597] Example 1: Leaving delivery unattended when the recipient is absent.

[1598] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[1599] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[1600] 3. Terminal: The intercom's microphone captures the visitor's speech, such as "Delivery. I've come to deliver your package," and sends it to the server.

[1601] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[1602] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if the package was left unattended when the recipient was absent.

[1603] 6. Server: Generates a response message saying, "We are currently out, please leave your package in front of the door," and sends it to the terminal.

[1604] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[1605] 8. Server: Notifies the user, "The delivery has arrived. The package has been left in front of your door."

[1606] 9. Server: The server uses the Emotion API to sense the user's stress level and adjusts the notification content to a softer tone.

[1607] Example of a prompt:

[1608] "If your doorbell rings and a visitor asks you to leave a package, how would you respond?"

[1609] Example 2: Solicitation / Sales Strategies

[1610] 1. Terminal: When the intercom rings, an infrared sensor detects the visitor and sends an event signal to the server.

[1611] 2. Server: The generative artificial intelligence outputs a voice message saying, "Who is this? What is your request? Which family member is this for?"

[1612] 3. Terminal: The intercom's microphone captures the visitor's statement, "We have an announcement about a new service," and sends it to the server.

[1613] 4. Server: Receives audio data and converts it to text using the Google Cloud Speech-to-Text API.

[1614] 5. Server: The text is analyzed using Hugging Face's NLP model to determine if it is a solicitation / sales pitch.

[1615] 6. Server: Generates a response message saying, "My husband is unable to assist you, so please excuse me," and sends it to the terminal.

[1616] 7. Terminal: Use the intercom speaker to give instructions to visitors.

[1617] 8. Server: Notifies the user that "a solicitation was received and declined."

[1618] 9. Server: Use the Emotion API to check the user's status and dynamically adjust the tone and content of notifications.

[1619] In this way, the system efficiently handles interactions with visitors while taking into account the user's emotional state, thereby improving user safety and convenience.

[1620] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1621] Step 1:

[1622] Terminal: The intercom's sensor detects a visitor. As input, the sensor obtains information about the visitor's presence. When the sensor is triggered, an event signal is generated and sent to the server. As output, the server recognizes the visitor's presence.

[1623] Step 2:

[1624] Server: The server receives an event signal and activates the generative artificial intelligence. It receives the event signal as input. The generative artificial intelligence generates initial dialogue questions. It generates questions such as "Who is this?", "What can I help you with?", and "Who in your family is this addressed to?" as text and converts them into speech using speech synthesis technology. It generates the voiced questions as output and sends them to the terminal.

[1625] Step 3:

[1626] Terminal: The intercom's microphone captures and records the visitor's response. It receives the visitor's voice response as input. The recorded data is sent to the server. It also sends the voice data to the server as output.

[1627] Step 4:

[1628] Server: The server converts the received audio data into text data using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). It receives audio data as input. The speech recognition engine analyzes the audio and generates text data. It generates text data as output.

[1629] Step 5:

[1630] Server: The converted text data is analyzed using natural language processing (e.g., Hugging Face's NLP model). It receives text data as input. The natural language processing analyzes the text content and classifies the visitor's request into multiple categories. It generates classified category information as output.

[1631] Step 6:

[1632] Server: Generates an appropriate automated response message based on the analysis results. It receives classified category information as input. The generative artificial intelligence generates a response message such as, "We are currently out, please leave your package at the front door." It sends the generated response message to the terminal as output.

[1633] Step 7:

[1634] Terminal: The terminal transmits the response message received from the server to the visitor through the intercom speaker. Input: Receives the response message. Plays the voice message using the intercom speaker. Output: Transmits the response message to the visitor.

[1635] Step 8:

[1636] Server: The generated response and information regarding the visitor's request are sent to the user's device using a notification method. The server receives the response and request information as input. It then notifies the user's device using a notification method (e.g., push notification or email). The notification is then delivered to the user as output.

[1637] Step 9:

[1638] Server: The server contains an emotion engine that recognizes the user's emotions from text and voice input from the user's device. It receives user input data as input. It analyzes the user's emotional state using the Emotion API and other tools. It generates user emotion information as output.

[1639] Step 10:

[1640] Server: The emotion engine adjusts notification content based on the user's emotional state. It receives user emotion information as input. It adjusts the notification content and tone according to the emotional state to generate an appropriate notification message. As output, it sends the adjusted notification message to the user's device.

[1641] Step 11:

[1642] User: The user receives a notification on a device such as a smartphone and generates additional instructions or reply messages via the server as needed. Input: Receives the user notification. Creates a custom message (e.g., "I'll be right back, please wait a moment") as needed and sends it to the server. Output: The reply message is transmitted to the visitor.

[1643] (Application Example 2)

[1644] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1645] Traditional intercom systems often required manual interaction with visitors, making quick and appropriate responses difficult, especially when users were away from home, working remotely, or with children home alone. Furthermore, the lack of appropriate notifications and responses tailored to the user's emotional state often led to inconvenience and stress. As a result, many users felt a lack of both safety and convenience.

[1646] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a sensor means for detecting visitors, a means for receiving notifications from the sensor means and activating a generative artificial intelligence to perform an initial dialogue, a means for the generative artificial intelligence to convert speech into text, a means for analyzing the text and classifying the visitor's request, a means for generating an automatic response message according to the classified request and transmitting it to the visitor, a means for notifying the request and response content to the user's communication terminal, and a means for recognizing the user's emotions and adjusting the tone and content of the notification and automatic response message. This enables quick and appropriate responses even when the user is absent or busy, and allows for flexible notifications and responses according to the user's emotional state.

[1647] "Sensor means" refers to a device or system used to detect visitors.

[1648] A "server means" is a device or system that receives notifications from sensor means, activates a generative artificial intelligence, and conducts an initial dialogue.

[1649] "Generative artificial intelligence" refers to artificial intelligence software that converts speech into text, analyzes that text, and generates appropriate response messages.

[1650] "Voice recognition means" refers to a device or system that converts a visitor's voice into text.

[1651] A "natural language processing device" is a device or system that analyzes text to classify the visitor's request.

[1652] "Response means" refers to a device or system that generates an automated response message according to the classified requirements and transmits it to the visitor.

[1653] "Notification means" refers to a device or system that notifies the user's communication terminal of the aforementioned request and response content.

[1654] "Emotion recognition means" refers to a device or system that recognizes the user's emotions and adjusts the tone and content of the notification and automated response messages.

[1655] A "communication terminal" is an electronic device, including smartphones and tablets, that a user uses to receive notifications.

[1656] The system for implementing this invention automatically performs a series of operations from visitor detection to response and notification, and further recognizes the user's emotions to adjust notifications and responses accordingly. The specific form of this system is described below.

[1657] Hardware configuration

[1658] This system uses the following hardware:

[1659] Sensory devices (e.g., intercoms, security cameras, door sensors)

[1660] Server configuration (e.g., cloud server)

[1661] User's communication device (e.g., smartphone, tablet)

[1662] Software to use

[1663] This system is equipped with the following software:

[1664] Generative artificial intelligence (e.g. GPT-3)

[1665] Speech recognition methods (e.g., Google Speech Recognition API)

[1666] Natural language processing tools (e.g., TextBlob library)

[1667] Emotion recognition methods (e.g., TextBlob's sentiment analysis function)

[1668] Data processing and data calculation

[1669] The server performs the following data processing and calculations.

[1670] 1. Visitor detection

[1671] Sensors detect the presence of visitors. Examples include intercoms and door sensors.

[1672] 2. Initiating the initial dialogue

[1673] The server receives a notification from a sensor and activates a generative artificial intelligence (AI) to initiate an initial conversation with the visitor, such as "Who are you?". This AI then uses speech recognition to convert the visitor's voice into text.

[1674] 3. Classification of Visitor's Purpose

[1675] Text data converted by generative artificial intelligence is analyzed using natural language processing techniques to appropriately classify the visitor's purpose.

[1676] 4. Generation and transmission of response messages

[1677] Based on the categorized request, an automated response message is generated and communicated to the visitor via the response system.

[1678] 5. Notification to the user's terminal

[1679] Simultaneously, the server sends information about the response and the visitor's purpose to the user's communication terminal using a notification mechanism.

[1680] 6. User emotion recognition and notification adjustment

[1681] Furthermore, using emotion recognition mechanisms, the system recognizes the user's emotions from their text or voice input and dynamically adjusts the tone of notification content and response messages.

[1682] Specific example

[1683] The following are specific examples.

[1684] Example 1: Delivery company's response

[1685] 1. The sensor detects the delivery person.

[1686] 2. The server activates a generative artificial intelligence and responds, "Who is this?"

[1687] 3. The delivery person responds, "I've come to deliver your package."

[1688] 4. The speech recognition means converts the speech into text, and the natural language processing means analyzes the text.

[1689] 5. The server generates a response message saying, "We are currently out, please leave it at the front door," and transmits it to the delivery person via the response system.

[1690] 6. The user's communication device is notified with the message, "Your delivery has arrived. The package has been left in front of your door."

[1691] Example 2: Detection of a suspicious person

[1692] 1. The sensor detects a suspicious person.

[1693] 2. The server activates a generative artificial intelligence and responds, "Who is it?", but the suspicious person remains silent.

[1694] 3. The server identifies the person as suspicious, generates a warning message, and transmits it via the response system.

[1695] 4. Simultaneously, a warning notification is sent to the user's communication terminal.

[1696] Example of a prompt

[1697] "Person says: I'm here to deliver a package. User is stressed. Generate a soft response."

[1698] In this way, it becomes possible to respond quickly and appropriately even when the user is away or busy, and to provide flexible notifications and responses that are tailored to their emotional state.

[1699] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1700] Step 1:

[1701] A sensor detects the presence of a visitor. The input is the sensor (e.g., intercom, security camera, door sensor), and the output is a notification signal of the detection event. This notification signal is sent to the server.

[1702] Step 2:

[1703] The server receives a notification from a sensor. The input is a notification signal, and the server activates a generative artificial intelligence to initiate an initial conversation. Specifically, it generates questions such as "Who are you?" for the visitor and outputs them as voice through a response system.

[1704] Step 3:

[1705] The visitor's response is captured by the intercom's microphone and recorded in real time as audio data. The input is the visitor's voice, and the output is the recorded audio data. The terminal sends this audio data to the server.

[1706] Step 4:

[1707] The server converts the received audio data into text data using speech recognition tools (e.g., Google Speech Recognition API). The input is audio data, and the output is text data. Specifically, the speech recognition engine analyzes the audio waveform data and converts it into the corresponding text.

[1708] Step 5:

[1709] The server analyzes the converted text data using natural language processing tools (e.g., the TextBlob library). The input is text data, and the output is the category of requests resulting from the analysis. In this process, the content of the text is analyzed, and its intent and requirements are classified into categories such as receiving packages when absent, refusing solicitations, responding during remote meetings, and how to handle situations when children are home alone.

[1710] Step 6:

[1711] The server generates an automated response message based on the analysis results. The input is the category of the request, and the output is the automated response message. Specifically, a generative artificial intelligence generates an appropriate message and conveys it to the visitor. For example, if the recipient is absent and the package is to be left at the door, the server will generate a message such as, "We are currently absent, please leave your package at the front door."

[1712] Step 7:

[1713] The server transmits a generated response message to the visitor via the response mechanism. The input is an automated response message, and the output is an audio message to the visitor. Specifically, it is transmitted to the visitor through the intercom speaker.

[1714] Step 8:

[1715] The server sends the generated response and information about the visitor's purpose to the user's communication terminal using a notification mechanism. The input is the response and purpose information, and the output is a notification message to the user's terminal. Specifically, it is displayed as a notification on the user's smartphone.

[1716] Step 9:

[1717] The server uses emotion recognition to recognize the user's emotions and adjusts the tone and content of notifications and automated response messages accordingly. The input is the user's emotion data (e.g., text or voice input), and the output is the adjusted notification and response message. For example, if the user is feeling stressed, the notification content might be adjusted to a softer tone, such as "Your package has been left at your front door."

[1718] Step 10:

[1719] The user generates additional instructions or follow-up messages as needed, which are then transmitted to the visitor via the server. The input is the user's follow-up message, and the output is a voice message to the visitor. For example, the user might send a custom message such as, "I'll be right back, please wait a moment."

[1720] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1721] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1722] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1723] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1724] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1725] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1726] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1727] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1728] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1729] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1730] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1731] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1732] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1733] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1734] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1735] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1736] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1737] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1738] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1739] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1740] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1741] The following is further disclosed regarding the embodiments described above.

[1742] (Claim 1)

[1743] A sensor means for detecting visitors,

[1744] A server means that receives notifications from the aforementioned sensor means, activates a generative artificial intelligence, and performs an initial dialogue.

[1745] The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text,

[1746] A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose,

[1747] A response means for generating an automated response message according to the classified request and communicating it to the visitor,

[1748] A notification means for notifying the user's terminal of the aforementioned request and response content,

[1749] A system that includes this.

[1750] (Claim 2)

[1751] The system according to claim 1, wherein the automated response message is selected from receiving packages when absent, refusing solicitations, responding during remote meetings, and responding when children are home alone.

[1752] (Claim 3)

[1753] The system according to claim 1, further comprising means for the server means to receive a reply message from the user's terminal and transmit it to the visitor.

[1754] "Example 1"

[1755] (Claim 1)

[1756] A sensor means for detecting visitors,

[1757] A server means that receives notifications from the aforementioned sensor means, activates a generative artificial intelligence, and performs an initial dialogue.

[1758] The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text,

[1759] A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose,

[1760] A response means for generating an automated response message according to the classified request and communicating it to the visitor,

[1761] A notification means for notifying the user's communication device of the aforementioned request and response content,

[1762] A system that includes this.

[1763] (Claim 2)

[1764] The system according to claim 1, wherein the automated response message is selected from receiving packages when absent, refusing solicitations, responding during remote meetings, and responding when children are home alone.

[1765] (Claim 3)

[1766] The system according to claim 1, further comprising means for the server means to receive a reply message from the user's communication device and transmit it to the visitor.

[1767] "Application Example 1"

[1768] (Claim 1)

[1769] A detection means for detecting visitors,

[1770] A computing means that receives a notification from the aforementioned detection means and activates a generative artificial intelligence to perform an initial dialogue,

[1771] The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text,

[1772] A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose,

[1773] A response means for generating an automated response message according to the classified request and communicating it to the visitor,

[1774] A notification means for notifying the user's terminal of the aforementioned request and response content,

[1775] A speech synthesis means that generates a voice response to a visitor using a speech synthesis engine,

[1776] A push notification method that sends visitor information to the user's device,

[1777] A system that includes this.

[1778] (Claim 2)

[1779] The system according to claim 1, wherein the automated response message is selected from receiving packages when absent, refusing solicitations, responding during remote meetings, and responding when children are home alone.

[1780] (Claim 3)

[1781] The system according to claim 1, further comprising means for receiving a response message from the user's terminal and transmitting it to the visitor.

[1782] "Example 2 of combining an emotion engine"

[1783] (Claim 1)

[1784] A sensor means for detecting visitors,

[1785] A computing means that receives notifications from the aforementioned sensor means, activates a generative artificial intelligence, and performs an initial dialogue.

[1786] The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text,

[1787] A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose,

[1788] A response means for generating an automated response message according to the classified request and communicating it to the visitor,

[1789] A notification means for notifying the user's terminal of the aforementioned request and response content,

[1790] An emotion engine that recognizes the user's emotions and adjusts notification content and responses accordingly,

[1791] A system that includes this.

[1792] (Claim 2)

[1793] The system according to claim 1, wherein the automated response message is selected from receiving packages when absent, refusing solicitations, responding during remote meetings, and responding when children are home alone.

[1794] (Claim 3)

[1795] The system according to claim 1, further comprising means for the computer means to receive a response message from the user's terminal and transmit it ...

Claims

1. A sensor means for detecting visitors, A server means that receives notifications from the aforementioned sensor means, activates a generative artificial intelligence, and performs an initial dialogue. The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text, A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose, A response means for generating an automated response message according to the classified request and communicating it to the visitor, A notification means for notifying the user's terminal of the aforementioned request and response content, A system that includes this.

2. The system according to claim 1, wherein the automated response message is selected from receiving packages when absent, refusing solicitations, responding during remote meetings, and responding when children are home alone.

3. The system according to claim 1, further comprising means for the server means to receive a reply message from the user's terminal and transmit it to the visitor.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A