system

An AI-powered intercom system addresses the anxiety of elderly residents by automatically recognizing and responding to visitors, ensuring safety and convenience through speech recognition and automated responses.

JP2026103466APending Publication Date: 2026-06-24SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-12
Publication Date
2026-06-24

AI Technical Summary

Technical Problem

Elderly people living alone face anxiety and stress from inappropriate visitors such as door-to-door sales and fraud, necessitating a means to safely and quickly confirm the purpose of visitors and respond automatically.

Method used

An intercom system with recording, speech recognition, natural language processing, evaluation, and automated response generation to identify and respond to visitor requirements, reducing psychological burden.

Benefits of technology

The system automates visitor interactions, enhancing safety and convenience by accurately identifying and responding to visitors, thereby reducing resident burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026103466000001_ABST
    Figure 2026103466000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 An acquisition means for acquiring voice information of a visitor, A recognition means for converting voice information into character information, A processing means for analyzing the requests of the visitor, An evaluation means for judging whether the request is appropriate, A generation means for generating a response to the visitor based on the evaluation result, A transmission means for conveying the generated response to the visitor, A notification means for notifying the response to the resident, A connection means for the resident to establish communication with the visitor, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Responding to visitors to the home can be a psychological and physical burden, especially for elderly people living alone. Specifically, it is a problem that they feel anxiety and stress towards inappropriate visitors such as door-to-door sales and fraud. In such a situation, there is a need for means for residents to safely and quickly confirm the purpose of the visitors and automatically respond as necessary.

Means for Solving the Problems

[0005] The present invention solves the above-mentioned problems by providing an intercom system that includes recording means for acquiring the voice of a visitor, speech recognition means for converting the acquired voice into text, natural language processing means for analyzing the text and determining the visitor's requirements, evaluation means for evaluating whether the requirements are permitted, transmission means for generating a response based on the evaluation and conveying it to the visitor, and text notification means and feedback processing means for residents. This enables residents to safely interact with visitors through speech recognition and automated responses, thereby reducing their psychological burden.

[0006] "Recording means" refers to a device or system for acquiring audio data of visitors.

[0007] "Speech recognition means" refers to technology for converting acquired speech data into text data.

[0008] "Natural language processing means" refers to algorithms or techniques for analyzing text data and understanding the requirements of visitors.

[0009] An "evaluation tool" is a system that includes criteria and processes for determining whether a visitor's requirements are appropriate.

[0010] The "response generation means" is a function that creates a response message for the visitor based on the evaluated results.

[0011] A "transmission means" is a mechanism for conveying the generated response message to the visitor.

[0012] A "notification method" is a system that presents converted text data to residents, allowing them to choose how to respond as needed.

[0013] A "feedback processing mechanism" is a function that receives feedback from residents and updates the criteria to improve the system's response. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is an AI-powered intercom system that responds to visitors. This system is designed to automate interactions with visitors, providing both security and convenience. The program and processing of this system are described below in natural language.

[0036] server

[0037] The server plays a central role in the system. First, the server receives audio data from visitors transmitted from terminals. This audio data is converted into text data using speech recognition means within the server. The converted text data is analyzed by natural language processing means to determine the visitor's requirements. The server uses evaluation means to determine whether the requirements are permitted or not. Based on the evaluation results, a response message for the visitor is created by response generation means. The server sends this message to the terminal, which then conveys it to the visitor.

[0038] Furthermore, the server creates a notification presenting the converted text to the resident. If the resident requires a response, the server manages the connection for video or voice calls and enables manual response. It also receives feedback from the resident and activates a feedback processing mechanism to update the system's response criteria.

[0039] terminal

[0040] The terminal is equipped with a recording mechanism to acquire visitor voice data. When a visitor presses the intercom, it transmits a message from the server via a speaker. The terminal transmits the response message generated by the server to the visitor via voice or text, and plays a role in establishing communication between the resident and the visitor as needed.

[0041] User (resident)

[0042] Users can check visitor information in real time by receiving system notifications. The AI ​​automatically responds to visitors who are not relevant, while users also have the option to respond manually when necessary. Users can easily communicate with visitors based on their own criteria. They can also improve the AI's response standards by providing feedback.

[0043] Thus, the intercom system of the present invention can improve the efficiency of responding to visitors and reduce the burden on residents by applying AI technology. Specifically, the ability to safely identify visitors and automatically reject unwanted visits is essential for the effective operation of this system.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The terminal detects when the visitor's call button is pressed and records the visitor's voice. The recorded audio data is immediately sent to the server.

[0047] Step 2:

[0048] The server receives audio data from the terminal and converts it into text data using speech recognition technology. This conversion utilizes a speech recognition algorithm.

[0049] Step 3:

[0050] The server analyzes the converted text data using natural language processing techniques to understand the visitor's requirements. This involves keyword extraction and contextual understanding.

[0051] Step 4:

[0052] The server uses evaluation tools to determine whether the requirements are appropriate. These evaluation criteria include pre-configured whitelists and blacklists.

[0053] Step 5:

[0054] Based on the evaluation results, the server uses a response generation mechanism to create an appropriate response message. For example, it tells authorized visitors "The resident is available to answer the door" and unauthorized visitors "We're sorry, but they're out."

[0055] Step 6:

[0056] The server sends the generated response message to the terminal. The terminal then communicates this message to the visitor via voice output or text display.

[0057] Step 7:

[0058] If the server determines that the request is permitted, it will send a notification to the resident. If the resident decides to respond, the server will establish a video or voice call connection.

[0059] Step 8:

[0060] The user receives a notification and reviews the visitor's details. If necessary, they can choose how to respond and proceed with the process of directly interacting with the visitor.

[0061] Step 9:

[0062] After interacting with a visitor, users can provide feedback to the server. Based on this feedback, the server updates its system response standards to improve the quality of future interactions.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] There is a need to streamline visitor reception, reduce the burden on residents, and ensure the safety of visitors by appropriately understanding their needs. However, conventional intercom systems only record visitors' voices, making it difficult to fully understand their intentions and respond automatically. Furthermore, residents are unable to quickly grasp the presence and intentions of visitors, leading to inefficient responses. An innovative intercom system is needed to solve these problems.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes an acquisition means for acquiring sound, a conversion means for converting sound into text, an analysis means for analyzing visitor information, a determination means for determining whether the information is appropriate, and a processing means for processing the information using a generative AI model. This makes it possible to automatically understand the visitor's intent, automatically generate an appropriate response, and respond to the visitor quickly and effectively.

[0068] "Acquisition means" refers to devices or processes that have the function of taking in external audio or data and transmitting it within the system.

[0069] "Conversion means" refers to the process or device that converts acquired audio data into text information.

[0070] "Analysis means" refers to processes and devices for interpreting and evaluating the visitor's intentions and requirements based on the converted textual information.

[0071] A "decision-making tool" is a device or process that evaluates whether the visitor's requirements are appropriate based on the analyzed information and determines the response strategy.

[0072] "Processing means" refers to processes and devices that utilize generative AI models to efficiently analyze information, generate responses, and understand context.

[0073] "Transmission means" refers to the functions or processes used to convey the generated response to the visitor.

[0074] "Means of establishing communication" refers to devices or processes that enable connections for direct communication between residents and visitors.

[0075] This invention relates to an embodiment of an innovative AI-powered intercom system. This system automates interactions with visitors, providing security and convenience. The system primarily consists of a server, terminals, and users.

[0076] The server is the core of the system and involves multiple mechanisms. First, the server receives audio data from visitors sent from terminals. This audio data is converted into text data using speech recognition technology such as Google® Cloud Speech-to-Text as a conversion mechanism. This text data is then processed using the generative AI model GPT-3® as an analysis mechanism to understand the visitor's requirements. Based on the analyzed information, a decision mechanism evaluates whether the requirements are appropriate. Based on the evaluation result, a response message is generated, and this response is processed by the generative AI model. Finally, the server sends this response to the terminal. The server also generates notifications for residents and, if necessary, enables direct communication by establishing communication between residents and visitors.

[0077] The terminal is a device that acquires visitor voices and is equipped with a microphone and recording device. When a visitor presses the intercom's call button, the terminal starts recording and sends the data to the server. When a response message is sent from the server, the terminal uses its speaker to communicate it to the visitor by voice or text. It also plays a role in establishing communication with residents.

[0078] Users (residents) can receive notifications from the server via devices such as smartphones and computers and check visitor information. If the automated response is inappropriate or manual intervention is required, users can directly interact with visitors through communication methods. The server also receives feedback from users and optimizes the system through the response generation process.

[0079] For example, if a visitor says "Hello, delivery," the server converts the audio into text data "Hello, delivery," and uses an analysis tool to identify that "the visit is for delivery." Next, it generates a response asking "Which company is the package from?" and transmits it to the visitor via the terminal.

[0080] An example of a prompt would be: "Tell me about the new AI intercom system. Explain how the system processes and responds when a visitor says, 'I'm a salesman. Do you have a moment?'"

[0081] In this way, the system aims to function as a fully automated system that can handle everything from acquiring voice data to generating responses and transmitting information to residents.

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The terminal acquires audio through its built-in microphone when a visitor presses the intercom button. This audio data is then input to the server. Specifically, when a visitor says, "Hello, delivery," their voice is recorded and sent to the server as a digital signal.

[0085] Step 2:

[0086] The server receives audio data sent from the terminal and converts the audio into text data using a conversion method. Here, audio data is used as input, and speech recognition technology such as Google Cloud Speech-to-Text is used to obtain text data as output. Specifically, the audio "Hello, delivery" is converted into the text "Hello, delivery".

[0087] Step 3:

[0088] The server analyzes the converted text data using an analysis method based on a generative AI model to determine the visitor's requirements. The text data is taken as input, analyzed using models such as OpenAI® GPT-3, and the result of interpreting the intent is obtained as output. Specifically, from the text "Hello, delivery," it determines that "the visitor is a delivery person."

[0089] Step 4:

[0090] Based on the judgment result, the server generates an appropriate response message for the visitor using a response generation mechanism. It takes the analyzed information as input and generates a response message as output. Specifically, if it is determined that the visitor's intention is "delivery," the response "Which company is the package from?" is generated.

[0091] Step 5:

[0092] The server sends a generated response message to the terminal, which then transmits it to the visitor using its speaker. The response from the server is used as input, and a message in voice or text format is sent as output. For example, the terminal might use the speaker to ask the visitor, "Which company is this package from?"

[0093] Step 6:

[0094] The server sends notifications about visitors to residents and manages the means of setting up communication between residents and visitors. It takes parsed visitor information as input and outputs notifications to residents and the establishment of communication. Specifically, a notification appears on the resident's smartphone stating, "A delivery person is here. Do you want to answer the door?"

[0095] Step 7:

[0096] Users can communicate directly with visitors and provide feedback to the server after the interaction is complete. The system receives the results of the direct communication and the feedback as input and outputs the results of improvements to the system's response standards. For example, if a user sends feedback stating that "the response should be faster," the system's response will be improved.

[0097] (Application Example 1)

[0098] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0099] In recent years, security concerns have increased, and there is a growing demand for more efficient and secure visitor handling. However, dealing with visitors when residents are absent or busy is inconvenient, and there are challenges in securely identifying visitors. Conventional systems have found it difficult to solve these problems simultaneously.

[0100] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0101] In this invention, the server includes acquisition means for acquiring voice information of visitors, recognition means for converting voice information into text information, processing means for analyzing the visitor's requests, generation means for generating a response to the visitor based on the evaluation results, transmission means for conveying the generated response to the visitor, notification means for notifying the resident of the response, connection means for the resident to establish communication with the visitor, and feedback means for updating the resident's judgment information based on the response generation means. As a result, visitor handling when the resident is away is automated, and the resident can check visitor information in real time and make appropriate decisions.

[0102] "Means of acquisition" refers to devices or technologies for collecting audio information from visitors.

[0103] "Recognition means" refers to a device or technology that converts collected audio information into textual information.

[0104] "Processing means" refers to a device or technology that analyzes visitor requests from textual information.

[0105] "Evaluation means" refers to a device or technology used to determine whether a request is appropriate or not.

[0106] "Generating means" refers to a device or technology that generates a response to a visitor based on the evaluation results.

[0107] "Means of communication" refers to devices or technologies used to convey the generated response to the visitor.

[0108] "Notification means" refers to a device or technology for informing residents of the content of a response.

[0109] "Connection means" refers to devices or technologies that enable residents to establish communication with visitors.

[0110] A "feedback mechanism" is a device or technology for updating response generation criteria based on residents' judgment information.

[0111] To implement this application, the system is constructed as follows: The server is equipped with microphones and recording devices to acquire visitor voice data, and the voice data acquired through these is converted into text data using a speech recognition service such as Google Cloud Speech-to-Text. This text data is analyzed using natural language processing software such as a BERT model or spaCy to determine the visitor's request.

[0112] The server uses an evaluation means to determine the validity of the visitor's request and a response generation means, for example using GPT-3 / 4, generates an appropriate response message for the visitor. This message is sent back to the visitor as voice or text via a transmission means.

[0113] The server uses a push notification service like Firebase to notify residents of its status in real time. This allows residents to receive visitor information through an application on their smartphone or tablet. Through their device, residents can also initiate video calls with visitors using protocols such as WebRTC. This allows residents to respond to visitors even when they are away or busy.

[0114] For example, if a delivery person arrives and the resident is not home, the server can automatically respond and instruct the delivery person on a drop-off location for the package. Furthermore, the resident can check the details of this interaction in real time, even when they are away from home.

[0115] An example of a prompt might be, "What is the visitor's request? If you are a delivery person, please answer 'delivery'." This allows the AI ​​model to generate an appropriate response and communicate it to the visitor.

[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0117] Step 1:

[0118] The server receives audio data from the visitor sent from the terminal. The input is audio data, which the server then processes using its recognition system.

[0119] Step 2:

[0120] The server converts the received audio data into text data using the Google Cloud Speech-to-Text service. The input is audio data, and the output is text data. This conversion is performed by analyzing the audio waveform and identifying the corresponding character information.

[0121] Step 3:

[0122] The server analyzes the converted text data using the BERT model and spaCy's natural language processing technology to interpret the visitor's requests. The input is text data, and the output is analyzed data regarding the visitor's requests. The analysis extracts the meaning of the text and is performed based on specific keywords and phrases.

[0123] Step 4:

[0124] The server evaluates the analytical data using an evaluation tool and determines whether the visitor's request should be permitted. The input is the analytical data, and the output is the decision result. This decision is made based on criteria set by the AI, referencing past data and current patterns.

[0125] Step 5:

[0126] The server uses GPT-3 / 4 to generate an appropriate response to the visitor based on the evaluation results. The input is the judgment result, and the output is the response message. The generation AI model generates natural conversational sentences using pre-configured prompt sentences.

[0127] Step 6:

[0128] The server transmits the generated response message to the terminal and sends the response to the visitor in voice or text. The output is the voice or text presented to the visitor. The means of transmission is via the network, and the message is presented in a format that is easy for the visitor to understand.

[0129] Step 7:

[0130] The server notifies residents of the response in real time. The input is the response message, and the output is the notification information sent to the residents. Residents receive the notifications on their smartphones via a push notification service such as Firebase.

[0131] Step 8:

[0132] Residents can initiate a video call with a visitor using the connection method via an interface on their device. Inputs are notification information and the resident's selection, while output is the video call connection status. Real-time communication is possible via the WebRTC protocol.

[0133] Step 9:

[0134] Users (residents) provide feedback on responses and system behavior using feedback mechanisms, and the server receives this feedback to update its response criteria. The input is user feedback, and the output is the updated decision criteria. The feedback is reflected in the AI's learning model and used to improve the accuracy of response generation.

[0135] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0136] This invention relates to an AI-powered intercom system incorporating an emotion engine that recognizes the emotions of visitors. The system aims to generate more appropriate responses by analyzing the visitor's voice and understanding their emotions. The embodiments of this invention are described in detail below.

[0137] server

[0138] In addition to its conventional functions, the server uses an emotion engine to analyze the emotions of visitors from their voices. When a visitor's voice is sent to the server, it first converts it into text data using speech recognition. The converted text data is then analyzed by natural language processing to understand the visitor's requirements. Subsequently, the server uses the emotion engine to perform emotion analysis and identify the visitor's emotional state (e.g., anger or joy).

[0139] The emotion analysis results are input into the response generation system, and the content and tone of the response are adjusted according to the emotions obtained. For example, if a visitor shows anxiety, it is possible to change the wording and tone of the voice, such as using reassuring language. The results of the emotion analysis are also reflected in the evaluation system, which helps to accurately understand the visitor's intentions.

[0140] terminal

[0141] The terminal continues to record the visitor's voice and send it to the server. Furthermore, the response message sent from the server is conveyed to the visitor. If necessary, operations are also performed to establish a call with the resident.

[0142] User (resident)

[0143] Users receive notifications from the server, allowing them to learn about visitors' information, including their emotional state. This emotional information enables users to better assess how to interact with visitors. For example, if a visitor shows signs of anxiety, they can take appropriate action, such as responding more carefully. Furthermore, users can communicate their own emotions to the server during feedback, which helps update response criteria.

[0144] Thus, by using an emotion engine to identify the emotions of visitors and optimizing the response, this invention realizes a safer and more meaningful intercom system for residents. For example, when a delivery person is about to deliver a package, if the resident is in a hurry, the system will respond accordingly, enabling flexible responses that reflect emotional information.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] The terminal detects when a visitor presses the intercom call button and immediately records the visitor's voice. The recorded voice data is sent to the server in real time.

[0148] Step 2:

[0149] The server first converts the audio data received from the terminal into text data using speech recognition technology. This allows the content of the audio to be obtained as text information.

[0150] Step 3:

[0151] The server analyzes the converted text data using natural language processing techniques to identify the requirements communicated by the visitor. This requirement identification includes keyword extraction and contextual analysis.

[0152] Step 4:

[0153] The server uses an emotion engine to recognize emotions from the visitor's voice data. Emotion recognition involves analyzing the tone and volume of the voice, as well as emotionally expressive words in the text.

[0154] Step 5:

[0155] The server combines the results of sentiment analysis with evaluation methods to determine whether the visitor's requirements are appropriate. The evaluation takes into account the visitor's emotional state.

[0156] Step 6:

[0157] Based on the evaluation results and sentiment recognition results, the server generates a response message using a response generation mechanism. For example, if the visitor is in a hurry, a quick response message is created.

[0158] Step 7:

[0159] The server sends the generated response message to the terminal. The terminal conveys this message to the visitor through audio output or display.

[0160] Step 8:

[0161] If the server deems it necessary, it will send a notification to the resident. This notification will include visitor details and emotional status.

[0162] Step 9:

[0163] The user receives a notification and decides how to respond based on the visitor's requirements and sentiment information. If an appropriate response is selected, the user begins communicating with the visitor via the server.

[0164] Step 10:

[0165] After completing all interactions, users can provide feedback on their responses and interactions with visitors. The server uses this feedback to update the criteria for the sentiment engine and response generation methods, improving future responses.

[0166] (Example 2)

[0167] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0168] Conventional intercom systems only used the visitor's voice information and did not take their emotional state into consideration, making it difficult for residents to accurately understand the visitor's intentions. This could lead to inappropriate responses to visitors and hinder the establishment of effective communication.

[0169] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0170] In this invention, the server includes acquisition means for acquiring voice information, acoustic recognition means for converting voice information into text information, and natural language processing means and sentiment analysis means for analyzing the visitor's requirements and emotions. This enables a comprehensive analysis including the visitor's emotions, allowing residents to respond to visitors more appropriately and effectively.

[0171] "Acquisition means" refers to a device or process that captures and records a visitor's voice information in digital format.

[0172] "Acoustic recognition means" refers to a technology or system that analyzes audio information and converts it into text information.

[0173] "Natural language processing means" refers to technologies or algorithms for analyzing visitors' intentions and requests from text information.

[0174] "Sentiment analysis methods" refer to processes or tools for identifying and analyzing emotional states from textual information.

[0175] A "response generation means" is a system or method for constructing an appropriate response for a visitor based on the analysis results.

[0176] "Means of communication" refers to the technology or device that delivers the generated response to the visitor.

[0177] A "feedback processing mechanism" is a process or function for receiving opinions and feedback from residents regarding their responses and for improving the system based on that information.

[0178] This invention realizes an AI-powered intercom system that analyzes emotions based on visitor voice information and provides appropriate responses. The system is broadly composed of a server, terminals, and users, each performing a specific function.

[0179] terminal

[0180] The terminal is a device that acquires audio information from visitors. Using a high-sensitivity microphone, the terminal records the visitor's voice as digital data and transmits this audio information to the server. By incorporating noise cancellation technology, it reduces external environmental noise and provides clear audio to the server.

[0181] server

[0182] The server is the core computer system that performs the processing. First, the server uses speech recognition software (e.g., a common speech recognition API) to convert the audio information sent from the terminal into text information. This text information is then analyzed using a natural language processing library (e.g., a common natural language processing API) to clarify the visitor's requirements and intentions. Next, sentiment analysis software (e.g., a common sentiment analysis API) analyzes the visitor's emotional state from the text information. This determines whether the visitor is expressing anger, joy, anxiety, etc. Finally, based on the analysis results, the server uses a response generation module to generate responses with different languages ​​and tones. These responses are adjusted to correspond to the visitor's emotions.

[0183] User (resident)

[0184] The user receives responses from the server and interacts with visitors. Received responses are converted into speech using speech synthesis technology and communicated to visitors through their device. Users can also send feedback to the server based on their interactions with visitors. This feedback is used to improve the server's response generation module.

[0185] For example, if a delivery person arrives to deliver a package and the system captures audio information indicating that the visitor is in a hurry, the server can generate a quick and efficient response. Through this embodiment, it becomes possible to understand the visitor's emotions and respond flexibly and effectively. An example of a prompt for effectively operating this system would be: "Analyze the visitor's audio data when the delivery person arrives to deliver the package, identify their emotions, and generate an appropriate response."

[0186] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0187] Step 1:

[0188] The device uses a highly sensitive microphone to capture the visitor's voice. The input is the visitor's voice, and the output is generated as digital audio data. This data is temporarily stored within the device and then prepared for transfer to the server. The device uses noise-canceling technology to remove background noise and improve audio clarity.

[0189] Step 2:

[0190] The server receives audio data transmitted from the terminal. The input is digital audio data, and the output is generated as text data. This conversion uses a speech recognition engine. Specifically, it analyzes sound wave data and generates strings based on the audio patterns. In this process, an appropriate language model is used to consider dialects and pronunciation variations.

[0191] Step 3:

[0192] The server analyzes the converted text data using a natural language processing library. The input is text data, and the output is structured data that shows the analyzed requirements and intentions. This process understands the visitor's intent by extracting keywords and analyzing grammatical patterns from the text. For example, if a visitor says "I'm in a hurry," the urgency is tagged as an important requirement.

[0193] Step 4:

[0194] The server inputs structured data into an emotion analysis engine to identify emotional states. The input is requirements analysis results, and the output is an emotion category (e.g., anger, joy, anxiety). The emotion analysis calculates positive and negative emotion scores to assess the visitor's mental state. Its function here is to quantify the emotional nuances of textual expressions.

[0195] Step 5:

[0196] The server constructs a response using a response generation module based on the obtained emotional state and requirements. The input is the emotional state and requirements, and the output is the adjusted response text. This process adjusts pre-prepared response templates to match the emotional state, changing the tone and wording. For example, it generates phrases indicating a quick response for a visitor in a hurry.

[0197] Step 6:

[0198] The terminal receives a response text sent from the server and applies speech synthesis technology to convey it to the visitor. The input is the response text, and the output is synthesized speech. This speech is output from the terminal to the visitor through a speaker. The terminal is required to make the speech clear and reproduce a more human-like tone.

[0199] Step 7:

[0200] Users (residents) provide feedback based on the interaction results and send it to the server. The input is the user's feedback, and the output is a dataset that contributes to improving the response generation algorithm. This feedback process allows the server to continuously improve the accuracy and appropriateness of its responses. Users evaluate whether the interaction with the visitor went smoothly and describe specific areas for improvement.

[0201] (Application Example 2)

[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0203] Conventional visitor response systems have difficulty taking visitors' emotions into account and fail to provide flexible responses appropriate to the visitor's situation. Furthermore, in situations where considering emotions is required for efficient and accurate responses, it is difficult to provide the most appropriate service quickly.

[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an acquisition means for acquiring visitor voice data, a conversion means for converting voice data into text data, and an emotion analysis means for analyzing the visitor's emotions. This enables flexible and accurate responses in accordance with the visitor's emotions.

[0205] "Acquisition means" refers to a device or system for accurately collecting visitor voice data.

[0206] "Conversion means" refers to a technology or process that converts collected audio data into text data.

[0207] "Analysis method" refers to a method for analyzing visitor requirements from converted text data.

[0208] "Emotional analysis methods" are technologies used to identify a visitor's emotional state from their voice and text data.

[0209] The "response generation means" is a function that generates an appropriate response to the visitor based on the analysis results and the sentiment analysis results.

[0210] "Communication means" refers to the technology or device used to transmit the generated response to the visitor.

[0211] "Notification method" refers to a method of notifying residents or responders of the analyzed information and presenting them with available response options.

[0212] "Update methods" refer to the process of continuously improving response standards based on feedback from those responding.

[0213] In the system of the present invention, the server acquires the visitor's voice data and converts it into text data. The converted text data is processed using natural language processing techniques to analyze the visitor's requirements and gain an understanding of them. Subsequently, sentiment analysis means are used to identify the visitor's emotional state. Based on the analysis results, response generation means generates an appropriate response tailored to the visitor. This response is communicated to the visitor via communication means.

[0214] The hardware consists of a microphone for recording audio and a server computer for processing. The software uses the "speech_recognition" library for speech recognition and the "transformers" library for sentiment analysis.

[0215] As a concrete example, the server detects the emotions of citizens visiting a citizen service counter, and if the emotions are analyzed as being anxious, it generates a message to provide reassurance in addition to providing regular information. In this way, interactions that enhance visitor satisfaction become possible.

[0216] Examples of prompts that utilize generative AI models are as follows:

[0217] "Identify the emotion from the text: {text indicating the relevant stress}"

[0218] This system will enable responses that better meet the needs of citizens, thereby improving the quality of administrative services.

[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0220] Step 1:

[0221] The device captures the visitor's voice and sends it to the server. The voice data is input, and a high-quality microphone is used to minimize noise during recording.

[0222] Step 2:

[0223] The server converts the acquired audio data into text data using speech recognition. In this step, the "speech_recognition" library is used to analyze the audio data and output the corresponding text data.

[0224] Step 3:

[0225] The server analyzes the converted text data using natural language processing tools to understand the visitor's requirements. In this process, the text data is input into a natural language processing model to generate requirements data that includes information related to the visitor's request.

[0226] Step 4:

[0227] The server uses sentiment analysis to identify the visitor's emotions from text data. The "transformers" library is used for sentiment analysis, taking text as input and outputting labels indicating emotions and their intensity.

[0228] Step 5:

[0229] The server generates the optimal response based on sentiment analysis results and requirements data. This response generation mechanism generates text that matches the visitor's emotions and requirements, creating a message that combines appropriate information and emotional responses for the visitor.

[0230] Step 6:

[0231] The server sends the generated response to the terminal via a communication method, informing the visitor. The terminal then provides the visitor with a response message in voice or text, completing the interaction.

[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0235] [Second Embodiment]

[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0248] This invention is an AI-powered intercom system that responds to visitors. This system is designed to automate interactions with visitors, providing both security and convenience. The program and its processes are described below in natural language.

[0249] server

[0250] The server plays a central role in the system. First, the server receives audio data from visitors transmitted from terminals. This audio data is converted into text data using speech recognition means within the server. The converted text data is analyzed by natural language processing means to determine the visitor's requirements. The server uses evaluation means to determine whether the requirements are permitted or not. Based on the evaluation results, a response message for the visitor is created by response generation means. The server sends this message to the terminal, which then conveys it to the visitor.

[0251] Furthermore, the server creates a notification presenting the converted text to the resident. If the resident requires a response, the server manages the connection for video or voice calls and enables manual response. It also receives feedback from the resident and activates a feedback processing mechanism to update the system's response criteria.

[0252] terminal

[0253] The terminal is equipped with a recording mechanism to acquire visitor voice data. When a visitor presses the intercom, it transmits a message from the server via a speaker. The terminal transmits the response message generated by the server to the visitor via voice or text, and plays a role in establishing communication between the resident and the visitor as needed.

[0254] User (resident)

[0255] Users can check visitor information in real time by receiving system notifications. The AI ​​automatically responds to visitors who are not relevant, while users also have the option to respond manually when necessary. Users can easily communicate with visitors based on their own criteria. They can also improve the AI's response standards by providing feedback.

[0256] Thus, the intercom system of the present invention can improve the efficiency of responding to visitors and reduce the burden on residents by applying AI technology. Specifically, the ability to safely identify visitors and automatically reject unwanted visits is essential for the effective operation of this system.

[0257] The following describes the processing flow.

[0258] Step 1:

[0259] The terminal detects when the visitor's call button is pressed and records the visitor's voice. The recorded voice data is immediately sent to the server.

[0260] Step 2:

[0261] The server receives audio data from the terminal and converts it into text data using speech recognition technology. This conversion utilizes a speech recognition algorithm.

[0262] Step 3:

[0263] The server analyzes the converted text data using natural language processing techniques to understand the visitor's requirements. This involves keyword extraction and contextual understanding.

[0264] Step 4:

[0265] The server uses evaluation tools to determine whether the requirements are appropriate. These evaluation criteria include pre-configured whitelists and blacklists.

[0266] Step 5:

[0267] Based on the evaluation results, the server uses a response generation mechanism to create an appropriate response message. For example, it tells authorized visitors "The resident is available to answer the door" and unauthorized visitors "We're sorry, but they are not home."

[0268] Step 6:

[0269] The server sends the generated response message to the terminal. The terminal then communicates this message to the visitor via voice output or text display.

[0270] Step 7:

[0271] If the server determines that the request is permitted, it will send a notification to the resident. If the resident decides to respond, the server will establish a video or voice call connection.

[0272] Step 8:

[0273] The user receives a notification and reviews the visitor's details. If necessary, they can choose how to respond and proceed with the process of directly interacting with the visitor.

[0274] Step 9:

[0275] After interacting with a visitor, users can provide feedback to the server. Based on this feedback, the server updates its system response standards to improve the quality of future interactions.

[0276] (Example 1)

[0277] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0278] There is a need to streamline visitor reception, reduce the burden on residents, and ensure the safety of visitors by appropriately understanding their needs. However, conventional intercom systems only record visitors' voices, making it difficult to fully understand their intentions and respond automatically. Furthermore, residents are unable to quickly grasp the presence and intentions of visitors, leading to inefficient responses. An innovative intercom system is needed to solve these problems.

[0279] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means respectively.

[0280] In this invention, the server includes an acquisition means for acquiring sound, a conversion means for converting the sound into characters, an analysis means for analyzing the information of the visitor, a determination means for determining whether the information is appropriate, and a processing means for processing the information using a generated AI model. Thereby, it becomes possible to automatically understand the intention of the visitor, automatically generate an appropriate response, and respond to the visitor quickly and effectively.

[0281] The "acquisition means" is a device or process having a function of capturing external voice and data and transmitting it into the system.

[0282] The "conversion means" is a process or device for converting the acquired voice data into character information.

[0283] The "analysis means" is a process or device for interpreting and evaluating the intention and requirements of the visitor based on the converted character information.

[0284] The "determination means" is a device or process for evaluating whether the requirements of the visitor are appropriate based on the analyzed information and determining the response policy.

[0285] The "processing means" is a process or device for analyzing information quickly and using the generated AI model to generate a response and understand the context.

[0286] The "transmission means" is a function or process for transmitting the generated response to the visitor.

[0287] The "communication establishment means" is a device or process for enabling a direct communication connection between the resident and the visitor.

[0288] This invention relates to an embodiment of an innovative AI-powered intercom system. This system automates interactions with visitors, providing security and convenience. The system primarily consists of a server, terminals, and users.

[0289] The server is the core of the system and involves multiple mechanisms. First, the server receives audio data from visitors sent from terminals. This audio data is converted into text data using speech recognition technology such as Google Cloud Speech-to-Text as the conversion mechanism. This text data is then processed using the generative AI model GPT-3 as the analysis mechanism to understand the visitor's requirements. Based on the analyzed information, the decision mechanism evaluates whether the requirements are appropriate. Based on the evaluation result, a response message is generated, and this response is processed by the generative AI model. Finally, the server sends this response to the terminal. The server also generates notifications for residents and enables direct communication by establishing communication between residents and visitors as needed.

[0290] The terminal is a device that acquires visitor voices and is equipped with a microphone and recording device. When a visitor presses the intercom's call button, the terminal starts recording and sends the data to the server. When a response message is sent from the server, the terminal uses its speaker to communicate it to the visitor by voice or text. It also plays a role in establishing communication with residents.

[0291] Users (residents) can receive notifications from the server via devices such as smartphones and computers and check visitor information. If the automated response is inappropriate or manual intervention is required, users can directly interact with visitors through communication methods. The server also receives feedback from users and optimizes the system through the response generation process.

[0292] For example, if a visitor says "Hello, delivery," the server converts the audio into text data "Hello, delivery," and uses an analysis tool to identify that "the visit is for delivery." Next, it generates a response asking "Which company is the package from?" and transmits it to the visitor via the terminal.

[0293] An example of a prompt would be: "Tell me about the new AI intercom system. Explain how the system processes and responds when a visitor says, 'I'm a salesman. Do you have a moment?'"

[0294] In this way, the system aims to function as a fully automated system that can handle everything from acquiring voice data to generating responses and transmitting information to residents.

[0295] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0296] Step 1:

[0297] The terminal acquires audio through its built-in microphone when a visitor presses the intercom button. This audio data is then input to the server. Specifically, when a visitor says, "Hello, delivery," their voice is recorded and sent to the server as a digital signal.

[0298] Step 2:

[0299] The server receives audio data sent from the terminal and converts the audio into text data using a conversion method. Here, audio data is used as input, and speech recognition technology such as Google Cloud Speech-to-Text is used to obtain text data as output. Specifically, the audio "Hello, delivery" is converted into the text "Hello, delivery".

[0300] Step 3:

[0301] The server analyzes the converted text data using an analysis means based on a generative AI model to determine the requirements of the visitor. Using the text data as input, it analyzes with a model such as OpenAI GPT-3 and obtains the result of interpreting the intention as output. As a specific operation, it determines the information that "the visitor is a delivery person" from the text "Hello, this is a delivery".

[0302] Step 4:

[0303] Based on the judgment result, the server uses a response generation means to generate an appropriate response message for the visitor. Using the analyzed information as input, it generates the response message as output. As a specific operation, when it is determined that the intention of the visitor is "delivery", the response "Which company's package is it?" is generated.

[0304] Step 5:

[0305] The server sends the generated response message to the terminal, and the terminal conveys it to the visitor using a speaker. Using the response from the server as input, it sends a message in voice or text form as output. As a specific operation, it conveys "Which company's package is it?" to the visitor through the speaker.

[0306] Step 6:

[0307] The server sends a notification regarding the visitor to the resident and manages the means for setting up communication between the resident and the visitor. Using the analyzed visitor information as input, it outputs the notification to the resident and the establishment of communication. As a specific operation, the notification "A delivery person has come. Will you respond?" is displayed on the resident's smartphone.

[0308] A Step 7:

[0309] Users can communicate directly with visitors and provide feedback to the server after the interaction is complete. The system receives the results of the direct communication and the feedback as input and outputs the results of improvements to the system's response standards. For example, if a user sends feedback stating that "the response should be faster," the system's response will be improved.

[0310] (Application Example 1)

[0311] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0312] In recent years, security concerns have increased, and there is a growing demand for more efficient and secure visitor handling. However, dealing with visitors when residents are absent or busy is inconvenient, and there are challenges in securely identifying visitors. Conventional systems have found it difficult to solve these problems simultaneously.

[0313] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0314] In this invention, the server includes acquisition means for acquiring voice information of visitors, recognition means for converting voice information into text information, processing means for analyzing the visitor's requests, generation means for generating a response to the visitor based on the evaluation results, transmission means for conveying the generated response to the visitor, notification means for notifying the resident of the response, connection means for the resident to establish communication with the visitor, and feedback means for updating the resident's judgment information based on the response generation means. As a result, visitor handling when the resident is away is automated, and the resident can check visitor information in real time and make appropriate decisions.

[0315] "Means of acquisition" refers to devices or technologies for collecting audio information from visitors.

[0316] "Recognition means" refers to a device or technology that converts collected audio information into textual information.

[0317] "Processing means" refers to a device or technology that analyzes visitor requests from textual information.

[0318] "Evaluation means" refers to a device or technology used to determine whether a request is appropriate or not.

[0319] "Generating means" refers to a device or technology that generates a response to a visitor based on the evaluation results.

[0320] "Means of communication" refers to devices or technologies used to convey the generated response to the visitor.

[0321] "Notification means" refers to a device or technology for informing residents of the content of a response.

[0322] "Connection means" refers to devices or technologies that enable residents to establish communication with visitors.

[0323] A "feedback mechanism" is a device or technology for updating response generation criteria based on residents' judgment information.

[0324] To implement this application, the system is constructed as follows: The server is equipped with microphones and recording devices to acquire visitor voice data, and the voice data acquired through these is converted into text data using a speech recognition service such as Google Cloud Speech-to-Text. This text data is analyzed using natural language processing software such as a BERT model or spaCy to determine the visitor's request.

[0325] The server uses an evaluation means to determine the validity of the visitor's request and a response generation means, for example using GPT-3 / 4, generates an appropriate response message for the visitor. This message is sent back to the visitor as voice or text via a transmission means.

[0326] The server uses a push notification service like Firebase to notify residents of its status in real time. This allows residents to receive visitor information through an application on their smartphone or tablet. Through their device, residents can also initiate video calls with visitors using protocols such as WebRTC. This allows residents to respond to visitors even when they are away or busy.

[0327] For example, if a delivery person arrives and the resident is not home, the server can automatically respond and instruct the delivery person on a drop-off location for the package. Furthermore, the resident can check the details of this interaction in real time, even when they are away from home.

[0328] An example of a prompt might be, "What is the visitor's request? If you are a delivery person, please answer 'delivery'." This allows the AI ​​model to generate an appropriate response and communicate it to the visitor.

[0329] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0330] Step 1:

[0331] The server receives audio data from the visitor sent from the terminal. The input is audio data, which the server then processes using its recognition system.

[0332] Step 2:

[0333] The server converts the received audio data into text data using the Google Cloud Speech-to-Text service. The input is audio data, and the output is text data. This conversion is performed by analyzing the audio waveform and identifying the corresponding character information.

[0334] Step 3:

[0335] The server analyzes the converted text data using the BERT model and spaCy's natural language processing technology to interpret the visitor's requests. The input is text data, and the output is analyzed data regarding the visitor's requests. The analysis extracts the meaning of the text and is performed based on specific keywords and phrases.

[0336] Step 4:

[0337] The server evaluates the analytical data using an evaluation tool and determines whether the visitor's request should be permitted. The input is the analytical data, and the output is the decision result. This decision is made based on criteria set by the AI, referencing past data and current patterns.

[0338] Step 5:

[0339] The server uses GPT-3 / 4 to generate an appropriate response to the visitor based on the evaluation results. The input is the judgment result, and the output is the response message. The generation AI model generates natural conversational sentences using pre-configured prompt sentences.

[0340] Step 6:

[0341] The server transmits the generated response message to the terminal and sends the response to the visitor in voice or text. The output is the voice or text presented to the visitor. The means of transmission is via the network, and the message is presented in a format that is easy for the visitor to understand.

[0342] Step 7:

[0343] The server notifies residents of the response in real time. The input is the response message, and the output is the notification information sent to the residents. Residents receive the notifications on their smartphones via a push notification service such as Firebase.

[0344] Step 8:

[0345] Residents can initiate a video call with a visitor using the connection method via an interface on their device. Inputs are notification information and the resident's selection, while output is the video call connection status. Real-time communication is possible via the WebRTC protocol.

[0346] Step 9:

[0347] Users (residents) provide feedback on responses and system behavior using feedback mechanisms, and the server receives this feedback to update its response criteria. The input is user feedback, and the output is the updated decision criteria. The feedback is reflected in the AI's learning model and used to improve the accuracy of response generation.

[0348] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0349] This invention relates to an AI-powered intercom system incorporating an emotion engine that recognizes the emotions of visitors. The system aims to generate more appropriate responses by analyzing the visitor's voice and understanding their emotions. The embodiments of this invention are described in detail below.

[0350] server

[0351] In addition to its conventional functions, the server uses an emotion engine to analyze the emotions of visitors from their voices. When a visitor's voice is sent to the server, it first converts it into text data using speech recognition. The converted text data is then analyzed by natural language processing to understand the visitor's requirements. Subsequently, the server uses the emotion engine to perform emotion analysis and identify the visitor's emotional state (e.g., anger or joy).

[0352] The emotion analysis results are input into the response generation system, and the content and tone of the response are adjusted according to the emotions obtained. For example, if a visitor shows anxiety, it is possible to change the wording and tone of the voice, such as using reassuring language. The results of the emotion analysis are also reflected in the evaluation system, which helps to accurately understand the visitor's intentions.

[0353] terminal

[0354] The terminal continues to record the visitor's voice and send it to the server. Furthermore, the response message sent from the server is conveyed to the visitor. If necessary, operations are also performed to establish a call with the resident.

[0355] User (resident)

[0356] Users receive notifications from the server, allowing them to learn about visitors' information, including their emotional state. This emotional information enables users to better assess how to interact with visitors. For example, if a visitor shows signs of anxiety, they can take appropriate action, such as responding more carefully. Furthermore, users can communicate their own emotions to the server during feedback, which helps update response criteria.

[0357] Thus, by using an emotion engine to identify the emotions of visitors and optimizing the response, this invention realizes a safer and more meaningful intercom system for residents. For example, when a delivery person is about to deliver a package, if the resident is in a hurry, the system will respond accordingly, enabling flexible responses that reflect emotional information.

[0358] The following describes the processing flow.

[0359] Step 1:

[0360] The terminal detects when a visitor presses the intercom call button and immediately records the visitor's voice. The recorded voice data is sent to the server in real time.

[0361] Step 2:

[0362] The server first converts the audio data received from the terminal into text data using speech recognition technology. This allows the content of the audio to be obtained as text information.

[0363] Step 3:

[0364] The server analyzes the converted text data using natural language processing techniques to identify the requirements communicated by the visitor. This requirement identification includes keyword extraction and contextual analysis.

[0365] Step 4:

[0366] The server uses an emotion engine to recognize emotions from the visitor's voice data. Emotion recognition involves analyzing the tone and volume of the voice, as well as emotionally expressive words in the text.

[0367] Step 5:

[0368] The server combines the results of sentiment analysis with evaluation methods to determine whether the visitor's requirements are appropriate. The evaluation takes into account the visitor's emotional state.

[0369] Step 6:

[0370] Based on the evaluation results and sentiment recognition results, the server generates a response message using a response generation mechanism. For example, if the visitor is in a hurry, a quick response message is created.

[0371] Step 7:

[0372] The server sends the generated response message to the terminal. The terminal conveys this message to the visitor through audio output or display.

[0373] Step 8:

[0374] If the server deems it necessary, it will send a notification to the resident. This notification will include visitor details and emotional status.

[0375] Step 9:

[0376] The user receives a notification and decides how to respond based on the visitor's requirements and sentiment information. If an appropriate response is selected, the user begins communicating with the visitor via the server.

[0377] Step 10:

[0378] After completing all interactions, users can provide feedback on their responses and interactions with visitors. The server uses this feedback to update the criteria for the sentiment engine and response generation methods, improving future responses.

[0379] (Example 2)

[0380] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0381] Conventional intercom systems only used the visitor's voice information and did not take their emotional state into consideration, making it difficult for residents to accurately understand the visitor's intentions. This could lead to inappropriate responses to visitors and hinder the establishment of effective communication.

[0382] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0383] In this invention, the server includes acquisition means for acquiring voice information, acoustic recognition means for converting voice information into text information, and natural language processing means and sentiment analysis means for analyzing the visitor's requirements and emotions. This enables a comprehensive analysis including the visitor's emotions, allowing residents to respond to visitors more appropriately and effectively.

[0384] "Acquisition means" refers to a device or process that captures and records a visitor's voice information in digital format.

[0385] "Acoustic recognition means" refers to a technology or system that analyzes audio information and converts it into text information.

[0386] "Natural language processing means" refers to technologies or algorithms for analyzing visitors' intentions and requests from text information.

[0387] "Sentiment analysis methods" refer to processes or tools for identifying and analyzing emotional states from textual information.

[0388] A "response generation means" is a system or method for constructing an appropriate response for a visitor based on the analysis results.

[0389] "Means of communication" refers to the technology or device that delivers the generated response to the visitor.

[0390] A "feedback processing mechanism" is a process or function for receiving opinions and feedback from residents regarding their responses and for improving the system based on that information.

[0391] This invention realizes an AI-powered intercom system that analyzes emotions based on visitor voice information and provides appropriate responses. The system is broadly composed of a server, terminals, and users, each performing a specific function.

[0392] terminal

[0393] The terminal is a device that acquires audio information from visitors. Using a high-sensitivity microphone, the terminal records the visitor's voice as digital data and transmits this audio information to the server. By incorporating noise cancellation technology, it reduces external environmental noise and provides clear audio to the server.

[0394] server

[0395] The server is the core computer system that performs the processing. First, the server uses speech recognition software (e.g., a common speech recognition API) to convert the audio information sent from the terminal into text information. This text information is then analyzed using a natural language processing library (e.g., a common natural language processing API) to clarify the visitor's requirements and intentions. Next, sentiment analysis software (e.g., a common sentiment analysis API) analyzes the visitor's emotional state from the text information. This determines whether the visitor is expressing anger, joy, anxiety, etc. Finally, based on the analysis results, the server uses a response generation module to generate responses with different languages ​​and tones. These responses are adjusted to correspond to the visitor's emotions.

[0396] User (resident)

[0397] The user receives responses from the server and interacts with visitors. Received responses are converted into speech using speech synthesis technology and communicated to visitors through their device. Users can also send feedback to the server based on their interactions with visitors. This feedback is used to improve the server's response generation module.

[0398] For example, if a delivery person arrives to deliver a package and the system captures audio information indicating that the visitor is in a hurry, the server can generate a quick and efficient response. Through this embodiment, it becomes possible to understand the visitor's emotions and respond flexibly and effectively. An example of a prompt for effectively operating this system would be: "Analyze the visitor's audio data when the delivery person arrives to deliver the package, identify their emotions, and generate an appropriate response."

[0399] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0400] Step 1:

[0401] The device uses a highly sensitive microphone to capture the visitor's voice. The input is the visitor's voice, and the output is generated as digital audio data. This data is temporarily stored within the device and then prepared for transfer to the server. The device uses noise-canceling technology to remove background noise and improve audio clarity.

[0402] Step 2:

[0403] The server receives audio data transmitted from the terminal. The input is digital audio data, and the output is generated as text data. This conversion uses a speech recognition engine. Specifically, it analyzes sound wave data and generates strings based on the audio patterns. In this process, an appropriate language model is used to consider dialects and pronunciation variations.

[0404] Step 3:

[0405] The server analyzes the converted text data using a natural language processing library. The input is text data, and the output is structured data that shows the analyzed requirements and intentions. This process understands the visitor's intent by extracting keywords and analyzing grammatical patterns from the text. For example, if a visitor says "I'm in a hurry," the urgency is tagged as an important requirement.

[0406] Step 4:

[0407] The server inputs structured data into an emotion analysis engine to identify emotional states. The input is requirements analysis results, and the output is an emotion category (e.g., anger, joy, anxiety). The emotion analysis calculates positive and negative emotion scores to assess the visitor's mental state. Its function here is to quantify the emotional nuances of textual expressions.

[0408] Step 5:

[0409] The server constructs a response using a response generation module based on the obtained emotional state and requirements. The input is the emotional state and requirements, and the output is the adjusted response text. This process adjusts pre-prepared response templates to match the emotional state, changing the tone and wording. For example, it generates phrases indicating a quick response for a visitor in a hurry.

[0410] Step 6:

[0411] The terminal receives a response text sent from the server and applies speech synthesis technology to convey it to the visitor. The input is the response text, and the output is synthesized speech. This speech is output from the terminal to the visitor through a speaker. The terminal is required to make the speech clear and reproduce a more human-like tone.

[0412] Step 7:

[0413] Users (residents) provide feedback based on the interaction results and send it to the server. The input is the user's feedback, and the output is a dataset that contributes to improving the response generation algorithm. This feedback process allows the server to continuously improve the accuracy and appropriateness of its responses. Users evaluate whether the interaction with the visitor went smoothly and describe specific areas for improvement.

[0414] (Application Example 2)

[0415] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0416] Conventional visitor response systems have difficulty taking visitors' emotions into account and fail to provide flexible responses appropriate to the visitor's situation. Furthermore, in situations where considering emotions is required for efficient and accurate responses, it is difficult to provide the most appropriate service quickly.

[0417] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an acquisition means for acquiring visitor voice data, a conversion means for converting voice data into text data, and an emotion analysis means for analyzing the visitor's emotions. This enables flexible and accurate responses in accordance with the visitor's emotions.

[0418] "Acquisition means" refers to a device or system for accurately collecting visitor voice data.

[0419] "Conversion means" refers to a technology or process that converts collected audio data into text data.

[0420] "Analysis method" refers to a method for analyzing visitor requirements from converted text data.

[0421] "Emotional analysis methods" are technologies used to identify a visitor's emotional state from their voice and text data.

[0422] The "response generation means" is a function that generates an appropriate response to the visitor based on the analysis results and the sentiment analysis results.

[0423] "Communication means" refers to the technology or device used to transmit the generated response to the visitor.

[0424] "Notification method" refers to a method of notifying residents or responders of the analyzed information and presenting them with available response options.

[0425] "Update methods" refer to the process of continuously improving response standards based on feedback from those responding.

[0426] In the system of the present invention, the server acquires the visitor's voice data and converts it into text data. The converted text data is processed using natural language processing techniques to analyze the visitor's requirements and gain an understanding of them. Subsequently, sentiment analysis means are used to identify the visitor's emotional state. Based on the analysis results, response generation means generates an appropriate response tailored to the visitor. This response is communicated to the visitor via communication means.

[0427] The hardware consists of a microphone for recording audio and a server computer for processing. The software uses the "speech_recognition" library for speech recognition and the "transformers" library for sentiment analysis.

[0428] As a concrete example, the server detects the emotions of citizens visiting a citizen service counter, and if the emotions are analyzed as being anxious, it generates a message to provide reassurance in addition to providing regular information. In this way, interactions that enhance visitor satisfaction become possible.

[0429] Examples of prompts that utilize generative AI models are as follows:

[0430] "Identify the emotion from the text: {text indicating the relevant stress}"

[0431] This system will enable responses that better meet the needs of citizens, thereby improving the quality of administrative services.

[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0433] Step 1:

[0434] The device captures the visitor's voice and sends it to the server. The voice data is input, and a high-quality microphone is used to minimize noise during recording.

[0435] Step 2:

[0436] The server converts the acquired audio data into text data using speech recognition. In this step, the "speech_recognition" library is used to analyze the audio data and output the corresponding text data.

[0437] Step 3:

[0438] The server analyzes the converted text data using natural language processing tools to understand the visitor's requirements. In this process, the text data is input into a natural language processing model to generate requirements data that includes information related to the visitor's request.

[0439] Step 4:

[0440] The server uses sentiment analysis to identify the visitor's emotions from text data. The "transformers" library is used for sentiment analysis, taking text as input and outputting labels indicating emotions and their intensity.

[0441] Step 5:

[0442] The server generates the optimal response based on sentiment analysis results and requirements data. This response generation mechanism generates text that matches the visitor's emotions and requirements, creating a message that combines appropriate information and emotional responses for the visitor.

[0443] Step 6:

[0444] The server sends the generated response to the terminal via a communication method, informing the visitor. The terminal then provides the visitor with a response message in voice or text, completing the interaction.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0448] [Third Embodiment]

[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0461] This invention is an AI-powered intercom system that responds to visitors. This system is designed to automate interactions with visitors, providing both security and convenience. The program and its processes are described below in natural language.

[0462] server

[0463] The server plays a central role in the system. First, the server receives audio data from visitors transmitted from terminals. This audio data is converted into text data using speech recognition means within the server. The converted text data is analyzed by natural language processing means to determine the visitor's requirements. The server uses evaluation means to determine whether the requirements are permitted or not. Based on the evaluation results, a response message for the visitor is created by response generation means. The server sends this message to the terminal, which then conveys it to the visitor.

[0464] Furthermore, the server creates a notification presenting the converted text to the resident. If the resident requires a response, the server manages the connection for video or voice calls and enables manual response. It also receives feedback from the resident and activates a feedback processing mechanism to update the system's response criteria.

[0465] terminal

[0466] The terminal is equipped with a recording mechanism to acquire visitor voice data. When a visitor presses the intercom, it transmits a message from the server via a speaker. The terminal transmits the response message generated by the server to the visitor via voice or text, and plays a role in establishing communication between the resident and the visitor as needed.

[0467] User (resident)

[0468] Users can check visitor information in real time by receiving system notifications. The AI ​​automatically responds to visitors who are not relevant, while users also have the option to respond manually when necessary. Users can easily communicate with visitors based on their own criteria. They can also improve the AI's response standards by providing feedback.

[0469] Thus, the intercom system of the present invention can improve the efficiency of responding to visitors and reduce the burden on residents by applying AI technology. Specifically, the ability to safely identify visitors and automatically reject unwanted visits is essential for the effective operation of this system.

[0470] The following describes the processing flow.

[0471] Step 1:

[0472] The terminal detects when the visitor's call button is pressed and records the visitor's voice. The recorded voice data is immediately sent to the server.

[0473] Step 2:

[0474] The server receives audio data from the terminal and converts it into text data using speech recognition technology. This conversion utilizes a speech recognition algorithm.

[0475] Step 3:

[0476] The server analyzes the converted text data using natural language processing techniques to understand the visitor's requirements. This involves keyword extraction and contextual understanding.

[0477] Step 4:

[0478] The server uses evaluation tools to determine whether the requirements are appropriate. These evaluation criteria include pre-configured whitelists and blacklists.

[0479] Step 5:

[0480] Based on the evaluation results, the server uses a response generation mechanism to create an appropriate response message. For example, it tells authorized visitors "The resident is available to answer the door" and unauthorized visitors "We're sorry, but they are not home."

[0481] Step 6:

[0482] The server sends the generated response message to the terminal. The terminal then communicates this message to the visitor via voice output or text display.

[0483] Step 7:

[0484] If the server determines that the request is permitted, it will send a notification to the resident. If the resident decides to respond, the server will establish a video or voice call connection.

[0485] Step 8:

[0486] The user receives a notification and reviews the visitor's details. If necessary, they can choose how to respond and proceed with the process of directly interacting with the visitor.

[0487] Step 9:

[0488] After interacting with a visitor, users can provide feedback to the server. Based on this feedback, the server updates its system response standards to improve the quality of future interactions.

[0489] (Example 1)

[0490] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0491] There is a need to streamline visitor reception, reduce the burden on residents, and ensure the safety of visitors by appropriately understanding their needs. However, conventional intercom systems only record visitors' voices, making it difficult to fully understand their intentions and respond automatically. Furthermore, residents are unable to quickly grasp the presence and intentions of visitors, leading to inefficient responses. An innovative intercom system is needed to solve these problems.

[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0493] In this invention, the server includes an acquisition means for acquiring sound, a conversion means for converting sound into text, an analysis means for analyzing visitor information, a determination means for determining whether the information is appropriate, and a processing means for processing the information using a generative AI model. This makes it possible to automatically understand the visitor's intent, automatically generate an appropriate response, and respond to the visitor quickly and effectively.

[0494] "Acquisition means" refers to devices or processes that have the function of taking in external audio or data and transmitting it into the system.

[0495] "Conversion means" refers to the process or device that converts acquired audio data into text information.

[0496] "Analysis means" refers to a process or device for interpreting and evaluating the visitor's intentions and requirements based on the converted textual information.

[0497] A "decision-making tool" is a device or process that evaluates whether the visitor's requirements are appropriate based on the analyzed information and determines the response strategy.

[0498] "Processing means" refers to processes and devices that utilize generative AI models to efficiently analyze information, generate responses, and understand context.

[0499] "Transmission means" refers to the functions or processes used to convey the generated response to the visitor.

[0500] "Means of establishing communication" refers to devices or processes that enable connections for direct communication between residents and visitors.

[0501] This invention relates to an embodiment of an innovative AI-powered intercom system. This system automates interactions with visitors, providing security and convenience. The system primarily consists of a server, terminals, and users.

[0502] The server is the core of the system and involves multiple mechanisms. First, the server receives audio data from visitors sent from terminals. This audio data is converted into text data using speech recognition technology such as Google Cloud Speech-to-Text as the conversion mechanism. This text data is then processed using the generative AI model GPT-3 as the analysis mechanism to understand the visitor's requirements. Based on the analyzed information, the decision mechanism evaluates whether the requirements are appropriate. Based on the evaluation result, a response message is generated, and this response is processed by the generative AI model. Finally, the server sends this response to the terminal. The server also generates notifications for residents and enables direct communication by establishing communication between residents and visitors as needed.

[0503] The terminal is a device that acquires visitor voices and is equipped with a microphone and recording device. When a visitor presses the intercom's call button, the terminal starts recording and sends the data to the server. When a response message is sent from the server, the terminal uses its speaker to communicate it to the visitor by voice or text. It also plays a role in establishing communication with residents.

[0504] Users (residents) can receive notifications from the server via devices such as smartphones and computers and check visitor information. If the automated response is inappropriate or manual intervention is required, users can directly interact with visitors through communication methods. The server also receives feedback from users and optimizes the system through the response generation process.

[0505] For example, if a visitor says "Hello, delivery," the server converts the audio into text data "Hello, delivery," and uses an analysis tool to identify that "the visit is for delivery." Next, it generates a response asking "Which company is the package from?" and transmits it to the visitor via the terminal.

[0506] An example of a prompt would be: "Tell me about the new AI intercom system. Explain how the system processes and responds when a visitor says, 'I'm a salesman. Do you have a moment?'"

[0507] In this way, the system aims to function as a fully automated system that can handle everything from acquiring voice data to generating responses and transmitting information to residents.

[0508] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0509] Step 1:

[0510] The terminal acquires audio through its built-in microphone when a visitor presses the intercom button. This audio data is then input to the server. Specifically, when a visitor says, "Hello, delivery," their voice is recorded and sent to the server as a digital signal.

[0511] Step 2:

[0512] The server receives audio data sent from the terminal and converts the audio into text data using a conversion method. Here, audio data is used as input, and speech recognition technology such as Google Cloud Speech-to-Text is used to obtain text data as output. Specifically, the audio "Hello, delivery" is converted into the text "Hello, delivery".

[0513] Step 3:

[0514] The server analyzes the converted text data using a generative AI model to determine the visitor's requirements. It takes text data as input, analyzes it using a model such as OpenAI GPT-3, and outputs the result of interpreting the intent. Specifically, it interprets the text "Hello, delivery" to determine that "the visitor is a delivery person."

[0515] Step 4:

[0516] Based on the judgment result, the server generates an appropriate response message for the visitor using a response generation mechanism. It takes the analyzed information as input and generates a response message as output. Specifically, if it is determined that the visitor's intention is "delivery," the response "Which company is the package from?" is generated.

[0517] Step 5:

[0518] The server sends a generated response message to the terminal, which then transmits it to the visitor using its speaker. The response from the server is used as input, and a message in voice or text format is sent as output. For example, the terminal might use the speaker to ask the visitor, "Which company is this package from?"

[0519] Step 6:

[0520] The server sends notifications about visitors to residents and manages the means of setting up communication between residents and visitors. It takes parsed visitor information as input and outputs notifications to residents and the establishment of communication. Specifically, a notification appears on the resident's smartphone stating, "A delivery person is here. Do you want to answer the door?"

[0521] Step 7:

[0522] Users can communicate directly with visitors and provide feedback to the server after the interaction is complete. The system receives the results of the direct communication and the feedback as input and outputs the results of improvements to the system's response standards. For example, if a user sends feedback stating that "the response should be faster," the system's response will be improved.

[0523] (Application Example 1)

[0524] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0525] In recent years, security concerns have increased, and there is a growing demand for more efficient and secure visitor handling. However, dealing with visitors when residents are absent or busy is inconvenient, and there are challenges in securely identifying visitors. Conventional systems have found it difficult to solve these problems simultaneously.

[0526] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0527] In this invention, the server includes acquisition means for acquiring voice information of visitors, recognition means for converting voice information into text information, processing means for analyzing the visitor's requests, generation means for generating a response to the visitor based on the evaluation results, transmission means for conveying the generated response to the visitor, notification means for notifying the resident of the response, connection means for the resident to establish communication with the visitor, and feedback means for updating the resident's judgment information based on the response generation means. As a result, visitor handling when the resident is away is automated, and the resident can check visitor information in real time and make appropriate decisions.

[0528] "Means of acquisition" refers to devices or technologies for collecting audio information from visitors.

[0529] "Recognition means" refers to a device or technology that converts collected audio information into textual information.

[0530] "Processing means" refers to a device or technology that analyzes visitor requests from textual information.

[0531] "Evaluation means" refers to a device or technology used to determine whether a request is appropriate or not.

[0532] "Generating means" refers to a device or technology that generates a response to a visitor based on the evaluation results.

[0533] "Means of communication" refers to devices or technologies used to convey the generated response to the visitor.

[0534] "Notification means" refers to a device or technology for informing residents of the content of a response.

[0535] "Connection means" refers to devices or technologies that enable residents to establish communication with visitors.

[0536] A "feedback mechanism" is a device or technology for updating response generation criteria based on residents' judgment information.

[0537] To implement this application, the system is constructed as follows: The server is equipped with microphones and recording devices to acquire visitor voice data, and the voice data acquired through these is converted into text data using a speech recognition service such as Google Cloud Speech-to-Text. This text data is analyzed using natural language processing software such as a BERT model or spaCy to determine the visitor's request.

[0538] The server uses an evaluation means to determine the validity of the visitor's request and a response generation means, for example using GPT-3 / 4, generates an appropriate response message for the visitor. This message is sent back to the visitor as voice or text via a transmission means.

[0539] The server uses a push notification service like Firebase to notify residents of its status in real time. This allows residents to receive visitor information through an application on their smartphone or tablet. Through their device, residents can also initiate video calls with visitors using protocols such as WebRTC. This allows residents to respond to visitors even when they are away or busy.

[0540] For example, if a delivery person arrives and the resident is not home, the server can automatically respond and instruct the delivery person on a drop-off location for the package. Furthermore, the resident can check the details of this interaction in real time, even when they are away from home.

[0541] An example of a prompt might be, "What is the visitor's request? If you are a delivery person, please answer 'delivery'." This allows the AI ​​model to generate an appropriate response and communicate it to the visitor.

[0542] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0543] Step 1:

[0544] The server receives audio data from the visitor sent from the terminal. The input is audio data, which the server then processes using its recognition system.

[0545] Step 2:

[0546] The server converts the received audio data into text data using the Google Cloud Speech-to-Text service. The input is audio data, and the output is text data. This conversion is performed by analyzing the audio waveform and identifying the corresponding character information.

[0547] Step 3:

[0548] The server analyzes the converted text data using the BERT model and spaCy's natural language processing technology to interpret the visitor's requests. The input is text data, and the output is analyzed data regarding the visitor's requests. The analysis extracts the meaning of the text and is performed based on specific keywords and phrases.

[0549] Step 4:

[0550] The server evaluates the analytical data using an evaluation tool and determines whether the visitor's request should be permitted. The input is the analytical data, and the output is the decision result. This decision is made based on criteria set by the AI, referencing past data and current patterns.

[0551] Step 5:

[0552] The server uses GPT-3 / 4 to generate an appropriate response to the visitor based on the evaluation results. The input is the judgment result, and the output is the response message. The generation AI model generates natural conversational sentences using pre-configured prompt sentences.

[0553] Step 6:

[0554] The server transmits the generated response message to the terminal and sends the response to the visitor in voice or text. The output is the voice or text presented to the visitor. The means of transmission is via the network, and the message is presented in a format that is easy for the visitor to understand.

[0555] Step 7:

[0556] The server notifies residents of the response in real time. The input is the response message, and the output is the notification information sent to the residents. Residents receive the notifications on their smartphones via a push notification service such as Firebase.

[0557] Step 8:

[0558] Residents can initiate video calls with visitors using the connection method via an interface on their device. Inputs are notification information and the resident's selection, while output is the video call connection status. Real-time communication is possible via the WebRTC protocol.

[0559] Step 9:

[0560] Users (residents) provide feedback on responses and system behavior using feedback mechanisms, and the server receives this feedback to update its response criteria. The input is user feedback, and the output is the updated decision criteria. The feedback is reflected in the AI's learning model and used to improve the accuracy of response generation.

[0561] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0562] This invention relates to an AI-powered intercom system incorporating an emotion engine that recognizes the emotions of visitors. The system aims to generate more appropriate responses by analyzing the visitor's voice and understanding their emotions. The embodiments of this invention are described in detail below.

[0563] server

[0564] In addition to its conventional functions, the server uses an emotion engine to analyze the emotions of visitors from their voices. When a visitor's voice is sent to the server, it first converts it into text data using speech recognition. The converted text data is then analyzed by natural language processing to understand the visitor's requirements. Subsequently, the server uses the emotion engine to perform emotion analysis and identify the visitor's emotional state (e.g., anger or joy).

[0565] The emotion analysis results are input into the response generation system, and the content and tone of the response are adjusted according to the emotions obtained. For example, if a visitor shows anxiety, it is possible to change the wording and tone of the voice, such as using reassuring language. The results of the emotion analysis are also reflected in the evaluation system, which helps to accurately understand the visitor's intentions.

[0566] terminal

[0567] The terminal continues to record the visitor's voice and send it to the server. Furthermore, the response message sent from the server is conveyed to the visitor. If necessary, operations are also performed to establish a call with the resident.

[0568] User (resident)

[0569] Users receive notifications from the server, allowing them to learn about visitors' information, including their emotional state. This emotional information enables users to better assess how to interact with visitors. For example, if a visitor shows signs of anxiety, they can take appropriate action, such as responding more carefully. Furthermore, users can communicate their own emotions to the server during feedback, which helps update response criteria.

[0570] Thus, by using an emotion engine to identify the emotions of visitors and optimizing the response, this invention realizes a safer and more meaningful intercom system for residents. For example, when a delivery person is about to deliver a package, if the resident is in a hurry, the system will respond accordingly, enabling flexible responses that reflect emotional information.

[0571] The following describes the processing flow.

[0572] Step 1:

[0573] The terminal detects when a visitor presses the intercom call button and immediately records the visitor's voice. The recorded voice data is sent to the server in real time.

[0574] Step 2:

[0575] The server first converts the audio data received from the terminal into text data using speech recognition technology. This allows the content of the audio to be obtained as text information.

[0576] Step 3:

[0577] The server analyzes the converted text data using natural language processing techniques to identify the requirements communicated by the visitor. This requirement identification includes keyword extraction and contextual analysis.

[0578] Step 4:

[0579] The server uses an emotion engine to recognize emotions from the visitor's voice data. Emotion recognition involves analyzing the tone and volume of the voice, as well as emotionally expressive words in the text.

[0580] Step 5:

[0581] The server combines the results of sentiment analysis with evaluation methods to determine whether the visitor's requirements are appropriate. The evaluation takes into account the visitor's emotional state.

[0582] Step 6:

[0583] Based on the evaluation results and sentiment recognition results, the server generates a response message using a response generation mechanism. For example, if the visitor is in a hurry, a quick response message is created.

[0584] Step 7:

[0585] The server sends the generated response message to the terminal. The terminal conveys this message to the visitor through audio output or display.

[0586] Step 8:

[0587] If the server deems it necessary, it will send a notification to the resident. This notification will include visitor details and emotional status.

[0588] Step 9:

[0589] The user receives a notification and decides how to respond based on the visitor's requirements and sentiment information. If an appropriate response is selected, the user begins communicating with the visitor via the server.

[0590] Step 10:

[0591] After completing all interactions, users can provide feedback on their responses and interactions with visitors. The server uses this feedback to update the criteria for the sentiment engine and response generation methods, improving future responses.

[0592] (Example 2)

[0593] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0594] Conventional intercom systems only used the visitor's voice information and did not take their emotional state into consideration, making it difficult for residents to accurately understand the visitor's intentions. This could lead to inappropriate responses to visitors and hinder the establishment of effective communication.

[0595] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0596] In this invention, the server includes acquisition means for acquiring voice information, acoustic recognition means for converting voice information into text information, and natural language processing means and sentiment analysis means for analyzing the visitor's requirements and emotions. This enables a comprehensive analysis including the visitor's emotions, allowing residents to respond to visitors more appropriately and effectively.

[0597] "Acquisition means" refers to a device or process that captures and records a visitor's voice information in digital format.

[0598] "Acoustic recognition means" refers to a technology or system that analyzes audio information and converts it into text information.

[0599] "Natural language processing means" refers to technologies or algorithms for analyzing visitors' intentions and requests from text information.

[0600] "Sentiment analysis methods" refer to processes or tools for identifying and analyzing emotional states from textual information.

[0601] A "response generation means" is a system or method for constructing an appropriate response for a visitor based on the analysis results.

[0602] "Means of communication" refers to the technology or device that delivers the generated response to the visitor.

[0603] A "feedback processing mechanism" is a process or function for receiving opinions and feedback from residents regarding their responses and for improving the system based on that information.

[0604] This invention realizes an AI-powered intercom system that analyzes emotions based on visitor voice information and provides appropriate responses. The system is broadly composed of a server, terminals, and users, each performing a specific function.

[0605] terminal

[0606] The terminal is a device that acquires audio information from visitors. Using a high-sensitivity microphone, the terminal records the visitor's voice as digital data and transmits this audio information to the server. By incorporating noise cancellation technology, it reduces external environmental noise and provides clear audio to the server.

[0607] server

[0608] The server is the core computer system that performs the processing. First, the server uses speech recognition software (e.g., a common speech recognition API) to convert the audio information sent from the terminal into text information. This text information is then analyzed using a natural language processing library (e.g., a common natural language processing API) to clarify the visitor's requirements and intentions. Next, sentiment analysis software (e.g., a common sentiment analysis API) analyzes the visitor's emotional state from the text information. This determines whether the visitor is expressing anger, joy, anxiety, etc. Finally, based on the analysis results, the server uses a response generation module to generate responses with different languages ​​and tones. These responses are adjusted to correspond to the visitor's emotions.

[0609] User (resident)

[0610] The user receives responses from the server and interacts with visitors. Received responses are converted into speech using speech synthesis technology and communicated to visitors through their device. Users can also send feedback to the server based on their interactions with visitors. This feedback is used to improve the server's response generation module.

[0611] For example, if a delivery person arrives to deliver a package and the system captures audio information indicating that the visitor is in a hurry, the server can generate a quick and efficient response. Through this embodiment, it becomes possible to understand the visitor's emotions and respond flexibly and effectively. An example of a prompt for effectively operating this system would be: "Analyze the visitor's audio data when the delivery person arrives to deliver the package, identify their emotions, and generate an appropriate response."

[0612] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0613] Step 1:

[0614] The device uses a highly sensitive microphone to capture the visitor's voice. The input is the visitor's voice, and the output is generated as digital audio data. This data is temporarily stored within the device and then prepared for transfer to the server. The device uses noise-canceling technology to remove background noise and improve audio clarity.

[0615] Step 2:

[0616] The server receives audio data transmitted from the terminal. The input is digital audio data, and the output is generated as text data. This conversion uses a speech recognition engine. Specifically, it analyzes sound wave data and generates strings based on the audio patterns. In this process, an appropriate language model is used to consider dialects and pronunciation variations.

[0617] Step 3:

[0618] The server analyzes the converted text data using a natural language processing library. The input is text data, and the output is structured data that shows the analyzed requirements and intentions. This process understands the visitor's intent by extracting keywords and analyzing grammatical patterns from the text. For example, if a visitor says "I'm in a hurry," the urgency is tagged as an important requirement.

[0619] Step 4:

[0620] The server inputs structured data into an emotion analysis engine to identify emotional states. The input is requirements analysis results, and the output is an emotion category (e.g., anger, joy, anxiety). The emotion analysis calculates positive and negative emotion scores to assess the visitor's mental state. Its function here is to quantify the emotional nuances of textual expressions.

[0621] Step 5:

[0622] The server constructs a response using a response generation module based on the obtained emotional state and requirements. The input is the emotional state and requirements, and the output is the adjusted response text. This process adjusts pre-prepared response templates to match the emotional state, changing the tone and wording. For example, it generates phrases indicating a quick response for a visitor in a hurry.

[0623] Step 6:

[0624] The terminal receives a response text sent from the server and applies speech synthesis technology to convey it to the visitor. The input is the response text, and the output is synthesized speech. This speech is output from the terminal to the visitor through a speaker. The terminal is required to make the speech clear and reproduce a more human-like tone.

[0625] Step 7:

[0626] Users (residents) provide feedback based on the interaction results and send it to the server. The input is the user's feedback, and the output is a dataset that contributes to improving the response generation algorithm. This feedback process allows the server to continuously improve the accuracy and appropriateness of its responses. Users evaluate whether the interaction with the visitor went smoothly and describe specific areas for improvement.

[0627] (Application Example 2)

[0628] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0629] Conventional visitor response systems have difficulty taking visitors' emotions into account and fail to provide flexible responses appropriate to the visitor's situation. Furthermore, in situations where considering emotions is required for efficient and accurate responses, it is difficult to provide the most appropriate service quickly.

[0630] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an acquisition means for acquiring visitor voice data, a conversion means for converting voice data into text data, and an emotion analysis means for analyzing the visitor's emotions. This enables flexible and accurate responses in accordance with the visitor's emotions.

[0631] "Acquisition means" refers to a device or system for accurately collecting visitor voice data.

[0632] "Conversion means" refers to a technology or process that converts collected audio data into text data.

[0633] "Analysis method" refers to a method for analyzing visitor requirements from converted text data.

[0634] "Emotional analysis methods" are technologies used to identify a visitor's emotional state from their voice and text data.

[0635] The "response generation means" is a function that generates an appropriate response to the visitor based on the analysis results and the sentiment analysis results.

[0636] "Communication means" refers to the technology or device used to transmit the generated response to the visitor.

[0637] "Notification method" refers to a method of notifying residents or responders of the analyzed information and presenting them with available response options.

[0638] "Update methods" refer to the process of continuously improving response standards based on feedback from those responding.

[0639] In the system of the present invention, the server acquires the visitor's voice data and converts it into text data. The converted text data is processed using natural language processing techniques to analyze the visitor's requirements and gain an understanding of them. Subsequently, sentiment analysis means are used to identify the visitor's emotional state. Based on the analysis results, response generation means generates an appropriate response tailored to the visitor. This response is communicated to the visitor via communication means.

[0640] The hardware consists of a microphone for recording audio and a server computer for processing. The software uses the "speech_recognition" library for speech recognition and the "transformers" library for sentiment analysis.

[0641] As a concrete example, the server detects the emotions of citizens visiting a citizen service counter, and if the emotions are analyzed as being anxious, it generates a message to provide reassurance in addition to providing regular information. In this way, interactions that enhance visitor satisfaction become possible.

[0642] Examples of prompts that utilize generative AI models are as follows:

[0643] "Identify the emotion from the text: {text indicating the relevant stress}"

[0644] This system will enable responses that better meet the needs of citizens, thereby improving the quality of administrative services.

[0645] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0646] Step 1:

[0647] The device captures the visitor's voice and sends it to the server. The voice data is input, and a high-quality microphone is used to minimize noise during recording.

[0648] Step 2:

[0649] The server converts the acquired audio data into text data using speech recognition. In this step, the "speech_recognition" library is used to analyze the audio data and output the corresponding text data.

[0650] Step 3:

[0651] The server analyzes the converted text data using natural language processing tools to understand the visitor's requirements. In this process, the text data is input into a natural language processing model to generate requirements data that includes information related to the visitor's request.

[0652] Step 4:

[0653] The server uses sentiment analysis to identify the visitor's emotions from text data. The "transformers" library is used for sentiment analysis, taking text as input and outputting labels indicating emotions and their intensity.

[0654] Step 5:

[0655] The server generates the optimal response based on sentiment analysis results and requirements data. This response generation mechanism generates text that matches the visitor's emotions and requirements, creating a message that combines appropriate information and emotional responses for the visitor.

[0656] Step 6:

[0657] The server sends the generated response to the terminal via a communication method, informing the visitor. The terminal then provides the visitor with a response message in voice or text, completing the interaction.

[0658] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0659] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0660] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0661] [Fourth Embodiment]

[0662] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0663] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0664] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0665] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0666] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0667] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0668] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0669] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0670] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0671] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0672] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0673] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0674] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0675] This invention is an AI-powered intercom system that responds to visitors. This system is designed to automate interactions with visitors, providing both security and convenience. The program and its processes are described below in natural language.

[0676] server

[0677] The server plays a central role in the system. First, the server receives audio data from visitors transmitted from terminals. This audio data is converted into text data using speech recognition means within the server. The converted text data is analyzed by natural language processing means to determine the visitor's requirements. The server uses evaluation means to determine whether the requirements are permitted or not. Based on the evaluation results, a response message for the visitor is created by response generation means. The server sends this message to the terminal, which then conveys it to the visitor.

[0678] Furthermore, the server creates a notification presenting the converted text to the resident. If the resident requires a response, the server manages the connection for video or voice calls and enables manual response. It also receives feedback from the resident and activates a feedback processing mechanism to update the system's response criteria.

[0679] terminal

[0680] The terminal is equipped with a recording mechanism to acquire visitor voice data. When a visitor presses the intercom, it transmits a message from the server via a speaker. The terminal transmits the response message generated by the server to the visitor via voice or text, and plays a role in establishing communication between the resident and the visitor as needed.

[0681] User (resident)

[0682] Users can check visitor information in real time by receiving system notifications. The AI ​​automatically responds to visitors who are not relevant, while users also have the option to respond manually when necessary. Users can easily communicate with visitors based on their own criteria. They can also improve the AI's response standards by providing feedback.

[0683] Thus, the intercom system of the present invention can improve the efficiency of responding to visitors and reduce the burden on residents by applying AI technology. Specifically, the ability to safely identify visitors and automatically reject unwanted visits is essential for the effective operation of this system.

[0684] The following describes the processing flow.

[0685] Step 1:

[0686] The terminal detects when the visitor's call button is pressed and records the visitor's voice. The recorded voice data is immediately sent to the server.

[0687] Step 2:

[0688] The server receives audio data from the terminal and converts it into text data using speech recognition technology. This conversion utilizes a speech recognition algorithm.

[0689] Step 3:

[0690] The server analyzes the converted text data using natural language processing techniques to understand the visitor's requirements. This involves keyword extraction and contextual understanding.

[0691] Step 4:

[0692] The server uses evaluation tools to determine whether the requirements are appropriate. These evaluation criteria include pre-configured whitelists and blacklists.

[0693] Step 5:

[0694] Based on the evaluation results, the server uses a response generation mechanism to create an appropriate response message. For example, it tells authorized visitors "The resident is available to answer the door" and unauthorized visitors "We're sorry, but they are not home."

[0695] Step 6:

[0696] The server sends the generated response message to the terminal. The terminal then communicates this message to the visitor via voice output or text display.

[0697] Step 7:

[0698] If the server determines that the request is permitted, it will send a notification to the resident. If the resident decides to respond, the server will establish a video or voice call connection.

[0699] Step 8:

[0700] The user receives a notification and reviews the visitor's details. If necessary, they can choose how to respond and proceed with the process of directly interacting with the visitor.

[0701] Step 9:

[0702] After interacting with a visitor, users can provide feedback to the server. Based on this feedback, the server updates its system response standards to improve the quality of future interactions.

[0703] (Example 1)

[0704] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0705] There is a need to streamline visitor reception, reduce the burden on residents, and ensure the safety of visitors by appropriately understanding their needs. However, conventional intercom systems only record visitors' voices, making it difficult to fully understand their intentions and respond automatically. Furthermore, residents are unable to quickly grasp the presence and intentions of visitors, leading to inefficient responses. An innovative intercom system is needed to solve these problems.

[0706] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0707] In this invention, the server includes an acquisition means for acquiring sound, a conversion means for converting sound into text, an analysis means for analyzing visitor information, a determination means for determining whether the information is appropriate, and a processing means for processing the information using a generative AI model. This makes it possible to automatically understand the visitor's intent, automatically generate an appropriate response, and respond to the visitor quickly and effectively.

[0708] "Acquisition means" refers to devices or processes that have the function of taking in external audio or data and transmitting it into the system.

[0709] "Conversion means" refers to the process or device that converts acquired audio data into text information.

[0710] "Analysis means" refers to a process or device for interpreting and evaluating the visitor's intentions and requirements based on the converted textual information.

[0711] A "decision-making tool" is a device or process that evaluates whether the visitor's requirements are appropriate based on the analyzed information and determines the response strategy.

[0712] "Processing means" refers to processes and devices that utilize generative AI models to efficiently analyze information, generate responses, and understand context.

[0713] "Transmission means" refers to the functions or processes used to convey the generated response to the visitor.

[0714] "Means of establishing communication" refers to devices or processes that enable connections for direct communication between residents and visitors.

[0715] This invention relates to an embodiment of an innovative AI-powered intercom system. This system automates interactions with visitors, providing security and convenience. The system primarily consists of a server, terminals, and users.

[0716] The server is the core of the system and involves multiple mechanisms. First, the server receives audio data from visitors sent from terminals. This audio data is converted into text data using speech recognition technology such as Google Cloud Speech-to-Text as the conversion mechanism. This text data is then processed using the generative AI model GPT-3 as the analysis mechanism to understand the visitor's requirements. Based on the analyzed information, the decision mechanism evaluates whether the requirements are appropriate. Based on the evaluation result, a response message is generated, and this response is processed by the generative AI model. Finally, the server sends this response to the terminal. The server also generates notifications for residents and enables direct communication by establishing communication between residents and visitors as needed.

[0717] The terminal is a device that acquires visitor voices and is equipped with a microphone and recording device. When a visitor presses the intercom's call button, the terminal starts recording and sends the data to the server. When a response message is sent from the server, the terminal uses its speaker to communicate it to the visitor by voice or text. It also plays a role in establishing communication with residents.

[0718] Users (residents) can receive notifications from the server via devices such as smartphones and computers and check visitor information. If the automated response is inappropriate or manual intervention is required, users can directly interact with visitors through communication methods. The server also receives feedback from users and optimizes the system through the response generation process.

[0719] For example, if a visitor says "Hello, delivery," the server converts the audio into text data "Hello, delivery," and uses an analysis tool to identify that "the visit is for delivery." Next, it generates a response asking "Which company is the package from?" and transmits it to the visitor via the terminal.

[0720] An example of a prompt would be: "Tell me about the new AI intercom system. Explain how the system processes and responds when a visitor says, 'I'm a salesman. Do you have a moment?'"

[0721] In this way, the system aims to function as a fully automated system that can handle everything from acquiring voice data to generating responses and transmitting information to residents.

[0722] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0723] Step 1:

[0724] The terminal acquires audio through its built-in microphone when a visitor presses the intercom button. This audio data is then input to the server. Specifically, when a visitor says, "Hello, delivery," their voice is recorded and sent to the server as a digital signal.

[0725] Step 2:

[0726] The server receives audio data sent from the terminal and converts the audio into text data using a conversion method. Here, audio data is used as input, and speech recognition technology such as Google Cloud Speech-to-Text is used to obtain text data as output. Specifically, the audio "Hello, delivery" is converted into the text "Hello, delivery".

[0727] Step 3:

[0728] The server analyzes the converted text data using a generative AI model to determine the visitor's requirements. It takes text data as input, analyzes it using a model such as OpenAI GPT-3, and outputs the result of interpreting the intent. Specifically, it interprets the text "Hello, delivery" to determine that "the visitor is a delivery person."

[0729] Step 4:

[0730] Based on the judgment result, the server generates an appropriate response message for the visitor using a response generation mechanism. It takes the analyzed information as input and generates a response message as output. Specifically, if it is determined that the visitor's intention is "delivery," the response "Which company is the package from?" is generated.

[0731] Step 5:

[0732] The server sends a generated response message to the terminal, which then transmits it to the visitor using its speaker. The response from the server is used as input, and a message in voice or text format is sent as output. For example, the terminal might use the speaker to ask the visitor, "Which company is this package from?"

[0733] Step 6:

[0734] The server sends notifications about visitors to residents and manages the means of setting up communication between residents and visitors. It takes parsed visitor information as input and outputs notifications to residents and the establishment of communication. Specifically, a notification appears on the resident's smartphone stating, "A delivery person is here. Do you want to answer the door?"

[0735] Step 7:

[0736] Users can communicate directly with visitors and provide feedback to the server after the interaction is complete. The system receives the results of the direct communication and the feedback as input and outputs the results of improvements to the system's response standards. For example, if a user sends feedback stating that "the response should be faster," the system's response will be improved.

[0737] (Application Example 1)

[0738] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0739] In recent years, security concerns have increased, and there is a growing demand for more efficient and secure visitor handling. However, dealing with visitors when residents are absent or busy is inconvenient, and there are challenges in securely identifying visitors. Conventional systems have found it difficult to solve these problems simultaneously.

[0740] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0741] In this invention, the server includes acquisition means for acquiring voice information of visitors, recognition means for converting voice information into text information, processing means for analyzing the visitor's requests, generation means for generating a response to the visitor based on the evaluation results, transmission means for conveying the generated response to the visitor, notification means for notifying the resident of the response, connection means for the resident to establish communication with the visitor, and feedback means for updating the resident's judgment information based on the response generation means. As a result, visitor handling when the resident is away is automated, and the resident can check visitor information in real time and make appropriate decisions.

[0742] "Means of acquisition" refers to devices or technologies for collecting audio information from visitors.

[0743] "Recognition means" refers to a device or technology that converts collected audio information into textual information.

[0744] "Processing means" refers to a device or technology that analyzes visitor requests from textual information.

[0745] "Evaluation means" refers to a device or technology used to determine whether a request is appropriate or not.

[0746] "Generating means" refers to a device or technology that generates a response to a visitor based on the evaluation results.

[0747] "Means of communication" refers to devices or technologies used to convey the generated response to the visitor.

[0748] "Notification means" refers to a device or technology for informing residents of the content of a response.

[0749] "Connection means" refers to devices or technologies that enable residents to establish communication with visitors.

[0750] A "feedback mechanism" is a device or technology for updating response generation criteria based on residents' judgment information.

[0751] To implement this application, the system is constructed as follows: The server is equipped with microphones and recording devices to acquire visitor voice data, and the voice data acquired through these is converted into text data using a speech recognition service such as Google Cloud Speech-to-Text. This text data is analyzed using natural language processing software such as a BERT model or spaCy to determine the visitor's request.

[0752] The server uses an evaluation means to determine the validity of the visitor's request and a response generation means, for example using GPT-3 / 4, generates an appropriate response message for the visitor. This message is sent back to the visitor as voice or text via a transmission means.

[0753] The server uses a push notification service like Firebase to notify residents of its status in real time. This allows residents to receive visitor information through an application on their smartphone or tablet. Through their device, residents can also initiate video calls with visitors using protocols such as WebRTC. This allows residents to respond to visitors even when they are away or busy.

[0754] For example, if a delivery person arrives and the resident is not home, the server can automatically respond and instruct the delivery person on a drop-off location for the package. Furthermore, the resident can check the details of this interaction in real time, even when they are away from home.

[0755] An example of a prompt might be, "What is the visitor's request? If you are a delivery person, please answer 'delivery'." This allows the AI ​​model to generate an appropriate response and communicate it to the visitor.

[0756] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0757] Step 1:

[0758] The server receives audio data from the visitor sent from the terminal. The input is audio data, which the server then processes using its recognition system.

[0759] Step 2:

[0760] The server converts the received audio data into text data using the Google Cloud Speech-to-Text service. The input is audio data, and the output is text data. This conversion is performed by analyzing the audio waveform and identifying the corresponding character information.

[0761] Step 3:

[0762] The server analyzes the converted text data using the BERT model and spaCy's natural language processing technology to interpret the visitor's requests. The input is text data, and the output is analyzed data regarding the visitor's requests. The analysis extracts the meaning of the text and is performed based on specific keywords and phrases.

[0763] Step 4:

[0764] The server evaluates the analytical data using an evaluation tool and determines whether the visitor's request should be permitted. The input is the analytical data, and the output is the decision result. This decision is made based on criteria set by the AI, referencing past data and current patterns.

[0765] Step 5:

[0766] The server uses GPT-3 / 4 to generate an appropriate response to the visitor based on the evaluation results. The input is the judgment result, and the output is the response message. The generation AI model generates natural conversational sentences using pre-configured prompt sentences.

[0767] Step 6:

[0768] The server transmits the generated response message to the terminal and sends the response to the visitor in voice or text. The output is the voice or text presented to the visitor. The means of transmission is via the network, and the message is presented in a format that is easy for the visitor to understand.

[0769] Step 7:

[0770] The server notifies residents of the response in real time. The input is the response message, and the output is the notification information sent to the residents. Residents receive the notifications on their smartphones via a push notification service such as Firebase.

[0771] Step 8:

[0772] Residents can initiate video calls with visitors using the connection method via an interface on their device. Inputs are notification information and the resident's selection, while output is the video call connection status. Real-time communication is possible via the WebRTC protocol.

[0773] Step 9:

[0774] Users (residents) provide feedback on responses and system behavior using feedback mechanisms, and the server receives this feedback to update its response criteria. The input is user feedback, and the output is the updated decision criteria. The feedback is reflected in the AI's learning model and used to improve the accuracy of response generation.

[0775] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0776] This invention relates to an AI-powered intercom system incorporating an emotion engine that recognizes the emotions of visitors. The system aims to generate more appropriate responses by analyzing the visitor's voice and understanding their emotions. The embodiments of this invention are described in detail below.

[0777] server

[0778] In addition to its conventional functions, the server uses an emotion engine to analyze the emotions of visitors from their voices. When a visitor's voice is sent to the server, it first converts it into text data using speech recognition. The converted text data is then analyzed by natural language processing to understand the visitor's requirements. Subsequently, the server uses the emotion engine to perform emotion analysis and identify the visitor's emotional state (e.g., anger or joy).

[0779] The emotion analysis results are input into the response generation system, and the content and tone of the response are adjusted according to the emotions obtained. For example, if a visitor shows anxiety, it is possible to change the wording and tone of the voice, such as using reassuring language. The results of the emotion analysis are also reflected in the evaluation system, which helps to accurately understand the visitor's intentions.

[0780] terminal

[0781] The terminal continues to record the visitor's voice and send it to the server. Furthermore, the response message sent from the server is conveyed to the visitor. If necessary, operations are also performed to establish a call with the resident.

[0782] User (resident)

[0783] Users receive notifications from the server, allowing them to learn about visitors' information, including their emotional state. This emotional information enables users to better assess how to interact with visitors. For example, if a visitor shows signs of anxiety, they can take appropriate action, such as responding more carefully. Furthermore, users can communicate their own emotions to the server during feedback, which helps update response criteria.

[0784] Thus, by using an emotion engine to identify the emotions of visitors and optimizing the response, this invention realizes a safer and more meaningful intercom system for residents. For example, when a delivery person is about to deliver a package, if the resident is in a hurry, the system will respond accordingly, enabling flexible responses that reflect emotional information.

[0785] The following describes the processing flow.

[0786] Step 1:

[0787] The terminal detects when a visitor presses the intercom call button and immediately records the visitor's voice. The recorded voice data is sent to the server in real time.

[0788] Step 2:

[0789] The server first converts the audio data received from the terminal into text data using speech recognition technology. This allows the content of the audio to be obtained as text information.

[0790] Step 3:

[0791] The server analyzes the converted text data using natural language processing techniques to identify the requirements communicated by the visitor. This requirement identification includes keyword extraction and contextual analysis.

[0792] Step 4:

[0793] The server uses an emotion engine to recognize emotions from the visitor's voice data. Emotion recognition involves analyzing the tone and volume of the voice, as well as emotionally expressive words in the text.

[0794] Step 5:

[0795] The server combines the results of sentiment analysis with evaluation methods to determine whether the visitor's requirements are appropriate. The evaluation takes into account the visitor's emotional state.

[0796] Step 6:

[0797] Based on the evaluation results and sentiment recognition results, the server generates a response message using a response generation mechanism. For example, if the visitor is in a hurry, a quick response message is created.

[0798] Step 7:

[0799] The server sends the generated response message to the terminal. The terminal conveys this message to the visitor through audio output or display.

[0800] Step 8:

[0801] If the server deems it necessary, it will send a notification to the resident. This notification will include visitor details and emotional status.

[0802] Step 9:

[0803] The user receives a notification and decides how to respond based on the visitor's requirements and sentiment information. If an appropriate response is selected, the user begins communicating with the visitor via the server.

[0804] Step 10:

[0805] After completing all interactions, users can provide feedback on their responses and interactions with visitors. The server uses this feedback to update the criteria for the sentiment engine and response generation methods, improving future responses.

[0806] (Example 2)

[0807] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0808] Conventional intercom systems only used the visitor's voice information and did not take their emotional state into consideration, making it difficult for residents to accurately understand the visitor's intentions. This could lead to inappropriate responses to visitors and hinder the establishment of effective communication.

[0809] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0810] In this invention, the server includes acquisition means for acquiring voice information, acoustic recognition means for converting voice information into text information, and natural language processing means and sentiment analysis means for analyzing the visitor's requirements and emotions. This enables a comprehensive analysis including the visitor's emotions, allowing residents to respond to visitors more appropriately and effectively.

[0811] "Acquisition means" refers to a device or process that captures and records a visitor's voice information in digital format.

[0812] "Acoustic recognition means" refers to a technology or system that analyzes audio information and converts it into text information.

[0813] "Natural language processing means" refers to technologies or algorithms for analyzing visitors' intentions and requests from text information.

[0814] "Sentiment analysis methods" refer to processes or tools for identifying and analyzing emotional states from textual information.

[0815] A "response generation means" is a system or method for constructing an appropriate response for a visitor based on the analysis results.

[0816] "Means of communication" refers to the technology or device that delivers the generated response to the visitor.

[0817] A "feedback processing mechanism" is a process or function for receiving opinions and feedback from residents regarding their responses and for improving the system based on that information.

[0818] This invention realizes an AI-powered intercom system that analyzes emotions based on visitor voice information and provides appropriate responses. The system is broadly composed of a server, terminals, and users, each performing a specific function.

[0819] terminal

[0820] The terminal is a device that acquires audio information from visitors. Using a high-sensitivity microphone, the terminal records the visitor's voice as digital data and transmits this audio information to the server. By incorporating noise cancellation technology, it reduces external environmental noise and provides clear audio to the server.

[0821] server

[0822] The server is the core computer system that performs the processing. First, the server uses speech recognition software (e.g., a common speech recognition API) to convert the audio information sent from the terminal into text information. This text information is then analyzed using a natural language processing library (e.g., a common natural language processing API) to clarify the visitor's requirements and intentions. Next, sentiment analysis software (e.g., a common sentiment analysis API) analyzes the visitor's emotional state from the text information. This determines whether the visitor is expressing anger, joy, anxiety, etc. Finally, based on the analysis results, the server uses a response generation module to generate responses with different languages ​​and tones. These responses are adjusted to correspond to the visitor's emotions.

[0823] User (resident)

[0824] The user receives responses from the server and interacts with visitors. Received responses are converted into speech using speech synthesis technology and communicated to visitors through their device. Users can also send feedback to the server based on their interactions with visitors. This feedback is used to improve the server's response generation module.

[0825] For example, if a delivery person arrives to deliver a package and the system captures audio information indicating that the visitor is in a hurry, the server can generate a quick and efficient response. Through this embodiment, it becomes possible to understand the visitor's emotions and respond flexibly and effectively. An example of a prompt for effectively operating this system would be: "Analyze the visitor's audio data when the delivery person arrives to deliver the package, identify their emotions, and generate an appropriate response."

[0826] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0827] Step 1:

[0828] The device uses a highly sensitive microphone to capture the visitor's voice. The input is the visitor's voice, and the output is generated as digital audio data. This data is temporarily stored within the device and then prepared for transfer to the server. The device uses noise-canceling technology to remove background noise and improve audio clarity.

[0829] Step 2:

[0830] The server receives audio data transmitted from the terminal. The input is digital audio data, and the output is generated as text data. This conversion uses a speech recognition engine. Specifically, it analyzes sound wave data and generates strings based on the audio patterns. In this process, an appropriate language model is used to consider dialects and pronunciation variations.

[0831] Step 3:

[0832] The server analyzes the converted text data using a natural language processing library. The input is text data, and the output is structured data that shows the analyzed requirements and intentions. This process understands the visitor's intent by extracting keywords and analyzing grammatical patterns from the text. For example, if a visitor says "I'm in a hurry," the urgency is tagged as an important requirement.

[0833] Step 4:

[0834] The server inputs structured data into an emotion analysis engine to identify emotional states. The input is requirements analysis results, and the output is an emotion category (e.g., anger, joy, anxiety). The emotion analysis calculates positive and negative emotion scores to assess the visitor's mental state. Its function here is to quantify the emotional nuances of textual expressions.

[0835] Step 5:

[0836] The server constructs a response using a response generation module based on the obtained emotional state and requirements. The input is the emotional state and requirements, and the output is the adjusted response text. This process adjusts pre-prepared response templates to match the emotional state, changing the tone and wording. For example, it generates phrases indicating a quick response for a visitor in a hurry.

[0837] Step 6:

[0838] The terminal receives a response text sent from the server and applies speech synthesis technology to convey it to the visitor. The input is the response text, and the output is synthesized speech. This speech is output from the terminal to the visitor through a speaker. The terminal is required to make the speech clear and reproduce a more human-like tone.

[0839] Step 7:

[0840] Users (residents) provide feedback based on the interaction results and send it to the server. The input is the user's feedback, and the output is a dataset that contributes to improving the response generation algorithm. This feedback process allows the server to continuously improve the accuracy and appropriateness of its responses. Users evaluate whether the interaction with the visitor went smoothly and describe specific areas for improvement.

[0841] (Application Example 2)

[0842] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0843] Conventional visitor response systems have difficulty taking visitors' emotions into account and fail to provide flexible responses appropriate to the visitor's situation. Furthermore, in situations where considering emotions is required for efficient and accurate responses, it is difficult to provide the most appropriate service quickly.

[0844] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an acquisition means for acquiring visitor voice data, a conversion means for converting voice data into text data, and an emotion analysis means for analyzing the visitor's emotions. This enables flexible and accurate responses in accordance with the visitor's emotions.

[0845] "Acquisition means" refers to a device or system for accurately collecting visitor voice data.

[0846] "Conversion means" refers to a technology or process that converts collected audio data into text data.

[0847] "Analysis method" refers to a method for analyzing visitor requirements from converted text data.

[0848] "Emotional analysis methods" are technologies used to identify a visitor's emotional state from their voice and text data.

[0849] The "response generation means" is a function that generates an appropriate response to the visitor based on the analysis results and the sentiment analysis results.

[0850] "Communication means" refers to the technology or device used to transmit the generated response to the visitor.

[0851] "Notification method" refers to a method of notifying residents or responders of the analyzed information and presenting them with available response options.

[0852] "Update methods" refer to the process of continuously improving response standards based on feedback from those responding.

[0853] In the system of the present invention, the server acquires the visitor's voice data and converts it into text data. The converted text data is processed using natural language processing techniques to analyze the visitor's requirements and gain an understanding of them. Subsequently, sentiment analysis means are used to identify the visitor's emotional state. Based on the analysis results, response generation means generates an appropriate response tailored to the visitor. This response is communicated to the visitor via communication means.

[0854] The hardware consists of a microphone for recording audio and a server computer for processing. The software uses the "speech_recognition" library for speech recognition and the "transformers" library for sentiment analysis.

[0855] As a concrete example, the server detects the emotions of citizens visiting a citizen service counter, and if the emotions are analyzed as being anxious, it generates a message to provide reassurance in addition to providing regular information. In this way, interactions that enhance visitor satisfaction become possible.

[0856] Examples of prompts that utilize generative AI models are as follows:

[0857] "Identify the emotion from the text: {text indicating the relevant stress}"

[0858] This system will enable responses that better meet the needs of citizens, thereby improving the quality of administrative services.

[0859] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0860] Step 1:

[0861] The device captures the visitor's voice and sends it to the server. The voice data is input, and a high-quality microphone is used to minimize noise during recording.

[0862] Step 2:

[0863] The server converts the acquired audio data into text data using speech recognition. In this step, the "speech_recognition" library is used to analyze the audio data and output the corresponding text data.

[0864] Step 3:

[0865] The server analyzes the converted text data using natural language processing tools to understand the visitor's requirements. In this process, the text data is input into a natural language processing model to generate requirements data that includes information related to the visitor's request.

[0866] Step 4:

[0867] The server uses sentiment analysis to identify the visitor's emotions from text data. The "transformers" library is used for sentiment analysis, taking text as input and outputting labels indicating emotions and their intensity.

[0868] Step 5:

[0869] The server generates the optimal response based on sentiment analysis results and requirements data. This response generation mechanism generates text that matches the visitor's emotions and requirements, creating a message that combines appropriate information and emotional responses for the visitor.

[0870] Step 6:

[0871] The server sends the generated response to the terminal via a communication method, informing the visitor. The terminal then provides the visitor with a response message in voice or text, completing the interaction.

[0872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0873] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0874] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0875] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0876] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0877] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0878] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0879] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0880] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0882] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0883] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0884] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0886] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0887] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0888] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0889] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0890] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0891] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0892] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0893] The following is further disclosed regarding the embodiments described above.

[0894] (Claim 1)

[0895] A recording means for acquiring audio data of visitors,

[0896] A speech recognition means for converting audio data into text data,

[0897] A natural language processing tool for analyzing visitor requirements,

[0898] An evaluation method for determining whether the requirements are appropriate,

[0899] A response generation means that generates a response to the visitor based on the evaluation results,

[0900] A system that includes a means of transmitting the generated response to the visitor.

[0901] (Claim 2)

[0902] The system according to claim 1, further comprising a notification means for presenting converted text data to a resident, allowing the resident to choose how to respond.

[0903] (Claim 3)

[0904] The system according to claim 1, further comprising a feedback processing means that receives feedback from residents and updates the judgment criteria of the response generation means.

[0905] "Example 1"

[0906] (Claim 1)

[0907] A means of acquiring sound,

[0908] A means of converting sound into text,

[0909] An analytical means for analyzing visitor information,

[0910] A means of determining whether the information is appropriate,

[0911] A generation means that generates a response based on the judgment result,

[0912] A means of sending the generated response to the visitor,

[0913] A means of establishing communication between residents and visitors,

[0914] A system that includes processing means for processing information using a generative AI model.

[0915] (Claim 2)

[0916] The system according to claim 1, which presents the converted text information to the resident and allows them to select a response.

[0917] (Claim 3)

[0918] The system according to claim 1, further comprising an update means for receiving resident responses and updating the criteria for generating responses.

[0919] "Application Example 1"

[0920] (Claim 1)

[0921] A means of acquiring visitor voice information,

[0922] A recognition means for converting audio information into text information,

[0923] A processing method for analyzing visitor requests,

[0924] A means of evaluation to determine whether the request is appropriate,

[0925] A generation means for generating responses to visitors based on evaluation results,

[0926] A means of communication that conveys the generated response to the visitor,

[0927] A means of notifying residents of the response,

[0928] A means of connection for residents to establish communication with visitors,

[0929] A system that includes this.

[0930] (Claim 2)

[0931] The system according to claim 1, further comprising a feedback means for updating resident judgment information based on a response generation means.

[0932] (Claim 3)

[0933] The system according to claim 1, comprising means for presenting the visitor's response status to the resident in real time, and enabling the resident to manually make an appropriate response.

[0934] "Example 2 of combining an emotion engine"

[0935] (Claim 1)

[0936] A means of acquiring visitor voice information,

[0937] A means for converting audio information into text information,

[0938] A natural language processing method and sentiment analysis method for analyzing the requirements and emotions of visitors,

[0939] A response generation means that generates and adjusts responses based on the visitor's emotional state,

[0940] A system that includes a means of communication for notifying visitors of the generated response.

[0941] (Claim 2)

[0942] The system according to claim 1, which presents converted text information and emotional states to the resident and allows the resident to choose how to respond.

[0943] (Claim 3)

[0944] The system according to claim 1, further comprising a feedback processing means for receiving feedback from residents and updating the criteria for the response generation means.

[0945] "Application example 2 when combining with an emotional engine"

[0946] (Claim 1)

[0947] A means of acquiring visitor voice data,

[0948] A conversion method for converting audio data into text data,

[0949] An analytical means for analyzing visitor requirements,

[0950] A means of analyzing visitors' emotions,

[0951] A response generation means that generates a response to a visitor based on the emotion analysis results,

[0952] A system that includes a means of communication to convey the generated response to the visitor.

[0953] (Claim 2)

[0954] The system according to claim 1, comprising a notification means for presenting converted text data and sentiment analysis results, enabling the responder to select a response.

[0955] (Claim 3)

[0956] The system according to claim 1, further comprising an update means for receiving feedback from the responder and updating the judgment criteria of the response generation means. [Explanation of Symbols]

[0957] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring visitor voice information, A recognition means for converting audio information into text information, A processing method for analyzing visitor requests, A means of evaluation to determine whether the request is appropriate, A generation means for generating responses to visitors based on evaluation results, A means of communication that conveys the generated response to the visitor, A means of notifying residents of the response, A means of connection for residents to establish communication with visitors, A system that includes this.

2. The system according to claim 1, further comprising a feedback means for updating resident judgment information based on a response generation means.

3. The system according to claim 1, comprising means for presenting the visitor's response status to the resident in real time, and enabling the resident to manually make an appropriate response.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A