System
The system addresses the burden and risks of direct visitor interaction by using sensors, voice analysis, and automated responses to enhance safety and comfort for vulnerable groups.
Patent Information
- Application Number
- JP2024118056
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Conventional intercom systems require direct, one-on-one interaction between visitors and users, placing a heavy burden on elderly people, children, and people with disabilities, and pose risks such as door-to-door sales, fraud, and robbery, necessitating a system that can respond appropriately without direct interaction and enhance safety.
A system that includes sensors to detect visitors, microphones to capture voice, software to analyze intent, speakers to output responses, digital signage to display QR codes, and a cloud server to record and analyze interactions, enabling automated and safe visitor management.
The system allows for efficient and safe interaction with visitors, reducing burdens on vulnerable groups and minimizing risks, while providing secure and comfortable living conditions.
Smart Images

Figure 2026017274000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional intercom systems require direct, one-on-one interaction between visitors and users, placing a heavy burden on elderly people, children, and people with disabilities, who have varying levels of judgment and physical strength. The system also poses a high risk of door-to-door salesmen, fraud, and robbery, which can sometimes cause anxiety and trouble. Furthermore, responding appropriately to visitors when users are away or busy requires time and effort, which can reduce the user's quality of life. A system that can solve these problems and support a safe and comfortable lifestyle was needed. [Means for solving the problem]
[0005] The present invention provides a system including a means for detecting the approach of a visitor, a means for capturing the visitor's voice, a means for analyzing the voice data to understand the visitor's intent, a means for generating a response based on the understood intent, a means for outputting the generated response as voice, a means for recording the interaction with the visitor, and a means for saving the recorded interaction. The response generation means also includes a means for displaying a QR code after understanding the visitor's intent, and a means for detecting suspicious behavior based on the recorded interaction. This allows users to respond appropriately to visitors without having to interact with them directly, reducing the burden on the elderly, children, and people with disabilities and reducing the risk of door-to-door sales, fraud, and robbery.
[0006] "Means for detecting approaching visitors" refers to devices such as sensors or cameras installed to detect when a visitor enters a specific area.
[0007] "Means for capturing visitor audio" refers to a microphone or other sound-gathering device that picks up and captures audio produced by a visitor as data.
[0008] "Means for analyzing voice data and understanding visitor intent" refers to software or a system that converts the voices made by visitors into digital data, analyzes the content of that data, and determines what the visitor is looking for.
[0009] "Means for generating a response based on the understood intent" refers to software or a system for automatically generating an appropriate response based on the visitor's intent obtained through speech analysis.
[0010] The "means for outputting the generated response by voice" refers to a speaker or other audio output device for uttering the automatically generated response content as voice.
[0011] "Means for recording visitor interactions" refers to devices or software for recording video and audio of contact between a visitor and the system.
[0012] "Means for storing the recorded response situation" refers to a storage device or cloud storage for storing recorded video and audio data.
[0013] "Means for displaying QR codes" refers to displays or digital signage that display QR codes so that visitors can make payments or make confirmations.
[0014] "Means for detecting suspicious behavior" refers to software or systems that analyze recorded response data and automatically detect abnormal behavior or signs of trouble. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to respond to visitors efficiently and safely. The following describes specific embodiments of the present invention.
[0037] System Configuration
[0038] The system consists of the following main components:
[0039] 1. Edge AI Camera: Detects visitor movements and captures footage.
[0040] 2. Audio pickup microphone: Captures the voice of visitors.
[0041] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[0042] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0043] 5. Speaker: Communicates the generated voice response to the visitor.
[0044] 6. QR code display device: Displays a QR code for payment and verification.
[0045] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[0046] Explanation of program processing
[0047] The system works by using an edge AI camera and a microphone to capture the visitor's voice and video when they approach and sending it to a cloud server.
[0048] Visitor detection and response
[0049] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0050] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[0051] Speech recognition and intent understanding
[0052] Device: A microphone captures the visitor's voice and sends it to the generative AI model.
[0053] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, "delivery."
[0054] Response generation and display
[0055] Server: Generates an appropriate response based on the information obtained from speech analysis, for example, "You're delivering a package. Please scan the QR code."
[0056] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[0057] QR code payment
[0058] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[0059] Terminal: Confirm that the payment has been completed and repeat the voice message, "Payment has been confirmed. Please place your luggage."
[0060] Recording and Monitoring
[0061] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[0062] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[0063] Specific examples
[0064] Examples:
[0065] This system will be installed on the doors of homes where elderly people live.
[0066] 1. Visitor detection and response
[0067] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[0068] 2. Speech Recognition and Intent Understanding
[0069] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[0070] 3. Response Generation and Display
[0071] The generative AI model generates a response saying, "This is a package delivery. Please scan the QR code," and the response is relayed to the visitor through a speaker. The QR code is then displayed on a digital sign.
[0072] 4. QR code payment
[0073] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[0074] 5. Recording and Monitoring
[0075] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[0076] The above is a specific embodiment for carrying out the present invention, which allows elderly people to safely and comfortably receive visitors and live with peace of mind even when they are away or busy.
[0077] The processing flow will be explained below.
[0078] Step 1:
[0079] Visitor Detection
[0080] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the detection area, the camera captures the visitor's video and transmits the information to the system.
[0081] Step 2:
[0082] Activating the microphone
[0083] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[0084] Step 3:
[0085] Digital human display and greeting
[0086] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0087] Step 4:
[0088] Capture audio data
[0089] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[0090] Step 5:
[0091] Analysis of audio data
[0092] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, such as "delivery."
[0093] Step 6:
[0094] Generating a response
[0095] Server: Generates an appropriate response based on the visitor's intent. For example, "You're about to deliver a package. Please scan the QR code."
[0096] Step 7:
[0097] Sending generated responses and audio output
[0098] Server: Generates and sends the response to the device.
[0099] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[0100] Step 8:
[0101] Displaying the QR code
[0102] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[0103] Step 9:
[0104] Payment confirmation
[0105] User: A visitor scans the QR code with their smartphone and makes a payment.
[0106] Terminal: If the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[0107] Step 10:
[0108] Record of response status
[0109] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[0110] Step 11:
[0111] Data storage and analysis
[0112] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[0113] Step 12:
[0114] User Notification
[0115] Server: Sends visitor information (images, audio, purpose, etc.) to the user and notifies them with a message such as "XX Takkyubin has been delivered."
[0116] Example 1
[0117] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0118] In modern society, responding to visitors needs to be done efficiently and safely. In particular, for elderly people and those living alone, face-to-face interactions with visitors are often a burden. Furthermore, when the home is busy or out of the home, prompt and accurate responses are required, and it is also important to detect suspicious behavior. To solve these issues, a system is needed that can automatically detect visitor movements, understand the visitor's intentions, and provide an appropriate response.
[0119] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0120] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intention, means for outputting the generated response as voice, means for displaying visual information to the visitor, means for recording the visitor's interaction, and means for saving the recorded interaction. This not only enables efficient interaction with visitors, but also enables the visitor's intention to be understood quickly and accurately, enabling appropriate interaction even when the visitor is absent or busy. Furthermore, by detecting suspicious behavior, safety is improved.
[0121] "Means for detecting approaching visitors" refers to devices such as sensors and cameras that detect visitor movement, as well as software that controls them.
[0122] "Means for capturing visitor audio" refers to a microphone for recording the audio made by the visitor, and any equipment or software for appropriately processing that audio data.
[0123] "Means for analyzing voice data to understand visitor intent" refers to algorithms and software for analyzing captured voice data and understanding its content, primarily using generative AI models.
[0124] "Means for generating responses based on understood intent" refers to software and algorithms that understand the visitor's intent and generate appropriate responses based on that content.
[0125] The "means for outputting the generated response as voice" refers to a speaker for outputting the generated response as voice, and a device and software for controlling the speaker.
[0126] "Means for displaying visual information to visitors" refers to devices such as displays and digital signage that provide visual information to visitors, as well as software for controlling them.
[0127] "Means for recording interactions with visitors" refers to devices such as cameras and microphones for recording and recording interactions with visitors and the situation, as well as software for controlling them.
[0128] "Means for storing recorded responses" refers to a storage device for safely storing video and audio data, and software for managing that data.
[0129] System configuration and operation overview
[0130] This system is a digital human-type AI reception system for automatically responding to visitors. The system consists of the following main hardware and software components:
[0131] 1. Edge AI Camera: Detects visitor movements and captures footage.
[0132] 2. Audio pickup microphone: Captures the voice of visitors.
[0133] 3. Digital Signage: Displaying a digital human and responding visually and audibly to visitors.
[0134] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0135] 5. Speaker: Communicates the generated voice response to the visitor.
[0136] 6. QR code display device: Displays a QR code for payment and verification.
[0137] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[0138] System operation details
[0139] Visitor Detection
[0140] Terminal: When the Edge AI camera detects a visitor's movement, the system automatically wakes up. At the same time, the audio pickup microphone activates and captures the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0141] Audio capture and transmission
[0142] Terminal: The visitor's voice (e.g., "This is a parcel delivery") is captured by a microphone and sent to a cloud server. Since the voice data is processed in real time, low-latency data transfer is required.
[0143] Voice analysis and intent understanding
[0144] Server: The cloud server analyzes the received voice data using a generative AI model to understand the visitor's intent. For example, it can extract the intent "delivery" from the voice data "This is a courier delivery."
[0145] Response generation and display
[0146] Server: The generative AI model generates an appropriate response based on the analysis results, for example, "This is a package delivery. Please scan the QR code."
[0147] Terminal: The text response sent from the server is communicated to the visitor using a digital sign and speaker. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the QR code is displayed on the digital sign.
[0148] QR code payment
[0149] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[0150] Terminal: Once the payment is confirmed, the visitor will be notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[0151] Recording and Monitoring
[0152] Terminal: Records and records interactions with visitors and sends the data to a cloud server.
[0153] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[0154] Example operation
[0155] Consider the case where this system is installed on the door of a home where elderly people live.
[0156] 1. Visitor Detection: When a courier approaches the door, the Edge AI camera detects the movement and the system is activated. A digital human appears and asks the courier, "Welcome. How can I help you?"
[0157] 2. Voice capture and transmission: The delivery person responds, "This is a courier delivery," and the voice is captured by the microphone and transmitted to the cloud server.
[0158] 3. Speech analysis and intent understanding: The generative AI model analyzes the voice data and understands the intent of "delivery."
[0159] 4. Response generation and display: The generative AI model generates a response such as "This is a package delivery. Please scan the QR code," plays the audio over the speaker, and displays the QR code on the digital signage.
[0160] 5. QR code payment: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, they will receive a voice message saying, "Payment has been confirmed. Please leave your package."
[0161] 6. Recording and monitoring: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[0162] As a result, this system can respond to visitors efficiently and safely, and can provide safe and secure living support for elderly people and those living alone in their homes.
[0163] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0164] Step 1: Visitor detection
[0165] Terminal: When the Edge AI camera detects a visitor's movement, the system wakes up. The audio pickup microphone activates and prepares to capture the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can we help you?"
[0166] Input: Visitor motion detection (Edge AI camera)
[0167] Output: System startup, digital human display, greeting message output
[0168] Step 2: Capture audio
[0169] Terminal: A microphone captures the visitor's voice, such as "This is a courier delivery," and the voice data is converted into a digital format.
[0170] Input: Visitor's voice
[0171] Output: Digital audio data
[0172] Step 3: Audio transmission and analysis
[0173] Terminal: The captured audio data is sent to the cloud server in real time.
[0174] Server: Analyzes the received voice data using a generative AI model to understand the visitor's intent. Converts the voice data into text and uses natural language processing technology to extract the intent. For example, understand the intent of "delivery" from the voice saying "This is a courier delivery."
[0175] Input: Digital audio data
[0176] Output: Text data containing intent
[0177] Step 4: Response Generation
[0178] Server: The generative AI model generates an appropriate response based on the analysis results, for example, a text response such as "This is a package delivery. Please scan the QR code."
[0179] Input: Text data containing visitor intent
[0180] Output: Response text
[0181] Step 5: Response display and audio output
[0182] Terminal: The response text received from the server is communicated to the visitor via a speaker and digital signage. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the digital signage displays the QR code.
[0183] Input: Response text
[0184] Output: Voice response, QR code display
[0185] Step 6: QR code payment
[0186] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[0187] Terminal: The QR code is scanned to confirm the payment has been completed. Once the payment is complete, the visitor is notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[0188] Input: QR code scan, payment information
[0189] Output: Payment completion notification
[0190] Step 7: Record and monitor
[0191] Terminal: Records and records interactions with visitors and sends the data, including video and audio data, to a cloud server.
[0192] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[0193] Input: Video data, audio data
[0194] Output: Saved data, suspicious behavior notification
[0195] (Application example 1)
[0196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0197] Modern stores are required to respond to visitors quickly and accurately, but this places a heavy burden on employees, making it difficult for visitors to receive satisfactory service. It is also important to detect suspicious behavior early and respond appropriately. The present invention aims to solve these problems and respond to visitors efficiently and safely. Another objective of the present invention is to enable employees to smoothly confirm visitor responses and improve the efficiency of store operations.
[0198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0199] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intent, means for generating a response based on the understood intent, means for outputting the generated response as voice, means for recording the interaction with the visitor, means for saving the recorded interaction, and means for an employee wearing smart glasses to visually confirm the response. This allows employees to instantly understand the visitor's intent through the smart glasses and respond promptly and appropriately. Furthermore, recording and saving the interaction makes it easy to review later and detect suspicious behavior.
[0200] A "means for detecting approaching visitors" is a device or method that senses the movement of a visitor and triggers activation of the system.
[0201] A "visitor voice capturing means" is a device or method that collects voice from a visitor and processes it as digital data.
[0202] The "means for analyzing voice data and understanding the visitor's intent" refers to a device or method for analyzing collected voice data and understanding the visitor's requests and questions.
[0203] The "means for generating a response based on the understood intent" is a device or method that automatically generates an appropriate response based on the analysis results.
[0204] The "means for outputting the generated response by voice" is a device or method for transmitting the generated response to the visitor as voice.
[0205] "Means for recording visitor interactions" refers to a device or method for saving the interaction and interaction with visitors as digital data.
[0206] "Means for storing recorded response situations" refers to a device or method for durably storing recorded digital data.
[0207] A "means for visually confirming a response by an employee equipped with smart glasses" is a device or method for visually confirming a response generated through smart glasses worn by an employee.
[0208] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and an embodiment thereof is shown based on an application example using smart glasses worn by employees.
[0209] System configuration:
[0210] The system consists of the following main components:
[0211] 1. Edge AI Camera: When a visitor enters the store, it detects their movement and activates the entire system.
[0212] 2. Audio pickup microphone: Collects visitors' voices and processes them as digital data.
[0213] 3. Smart glasses: Devices worn by employees to visually and audibly confirm the visitor's intentions.
[0214] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0215] 5. Speaker: Outputs the generated response as audio and conveys it to the visitor.
[0216] 6. Cloud server: Stores recorded responses and detects suspicious behavior.
[0217] Explanation of program operation:
[0218] The server controls the edge AI camera, microphone, smart glasses, generative AI model, speaker, and cloud server to enable interaction with visitors. The main processing flow is as follows:
[0219] Visitor Detection:
[0220] When the Edge AI camera detects a visitor's movement, the system is activated, the microphone captures the visitor's voice, and a digital human appears in the smart glasses and asks the visitor, "Welcome. How can I help you?"
[0221] Speech Recognition and Intent Understanding:
[0222] The system captures audio data with a microphone and sends it to a generative AI model, which then analyzes the audio data to understand the visitor's intent.
[0223] For example, if a visitor says, "Please tell me where the new products are," the generative AI model understands the intent "new products" and generates the response, "The new products corner is in the back right."
[0224] Response generation and display:
[0225] The generative AI model generates an appropriate response, which is output through the speaker and also displayed on the smart glasses' display.
[0226] Record of response:
[0227] Visitor interactions are videotaped and stored on a cloud server, and if any suspicious activity is detected, a notification is sent to the administrator.
[0228] Hardware and software used:
[0229] Hardware: Edge AI camera (general security camera), sound pickup microphone (standalone microphone), smart glasses (e.g., smart glasses), speaker.
[0230] Software: speech recognition models (e.g., Google Speech-to-Text API), generative AI models (e.g., OpenAI GPT), and cloud storage (e.g., Google Cloud).
[0231] Example prompt sentence:
[0232] Input speech data: "Please tell me where the new products are located."
[0233] Prompt: "Based on the audio data, analyze the visitor's intent and generate an appropriate response. For example, a response to a question about the location of a new product."
[0234] Thus, the present invention utilizes cutting edge technology such as smart glasses to provide a system that can efficiently and safely serve visitors.
[0235] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0236] Step 1:
[0237] Visitor Detection
[0238] The device's edge AI camera detects the visitor's movements. The motion detection activates the system, which then captures the visitor's video data. The input is the visitor's movements, and the output is the video data and a signal to activate the system.
[0239] Step 2:
[0240] Audio Capture
[0241] The device's microphone captures the visitor's voice, and the collected voice data is sent to the system. The input is the visitor's voice, and the output is digital voice data.
[0242] Step 3:
[0243] Analysis of audio data
[0244] The server's generated AI model analyzes the collected voice data and understands the visitor's intent. The input is digital voice data, and the output is the visitor's intent as an analysis result. A voice recognition model is used for the analysis.
[0245] Step 4:
[0246] Response Generation
[0247] The server's generative AI model generates an appropriate response based on the visitor's intent. The input is the visitor's intent as a result of analysis, and the output is the generated response text. For example, if the question is about the location of new products, the generated response will be "The new products corner is in the back right."
[0248] Step 5:
[0249] Display and speak responses
[0250] The smart glasses on the terminal visually display the generated response, and the speaker outputs the response aloud. The input is the generated response text, and the output is the visual display and audio output. Employees can respond to visitors' questions instantly through the smart glasses.
[0251] Step 6:
[0252] Record of response status
[0253] The device records the conversation with the visitor and sends it to the cloud server. The input is the audio and video data of the conversation with the visitor, and the output is the recorded data sent to the cloud server.
[0254] Step 7:
[0255] Data storage and analysis
[0256] The cloud server stores the recorded response data and analyzes suspicious behavior. The input is video and audio data, and the output is the analysis result, detecting suspicious behavior and notifying the administrator.
[0257] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0258] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to efficiently and safely serve visitors while recognizing their emotions and responding appropriately. A specific embodiment of the present invention will be described below.
[0259] System Configuration
[0260] The system consists of the following main components:
[0261] 1. Edge AI Camera: Detects visitor movements and captures footage.
[0262] 2. Audio pickup microphone: Captures the voice of visitors.
[0263] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[0264] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0265] 5. Emotion Engine: Analyzes audio and video data to recognize visitors' emotions.
[0266] 6. Speaker: Communicates the generated voice response to the visitor.
[0267] 7. QR code display device: Displays a QR code for payment and verification.
[0268] 8. Cloud server: Stores recorded data and analyzes suspicious behavior.
[0269] Explanation of program processing
[0270] The system works by capturing the visitor's voice and video using an edge AI camera and microphone when the visitor approaches, sending the captured video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[0271] Visitor detection and response
[0272] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0273] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[0274] Speech recognition and intent understanding
[0275] Device: A microphone captures the visitor's voice and sends it to a cloud server running a generative AI model.
[0276] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, determining the intent to "deliver."
[0277] emotion recognition
[0278] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone.
[0279] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice, such as "happy," "angry," or "sad."
[0280] Response generation and display
[0281] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[0282] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[0283] QR code payment
[0284] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[0285] Terminal: Confirms that the payment has been completed and repeats the message "Payment has been confirmed. Please place your luggage."
[0286] Recording and Monitoring
[0287] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[0288] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[0289] Specific examples
[0290] Examples:
[0291] This system will be installed on the doors of homes where elderly people live.
[0292] 1. Visitor detection and response
[0293] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[0294] 2. Speech Recognition and Intent Understanding
[0295] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[0296] 3. Emotion recognition
[0297] The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[0298] 4. Response Generation and Display
[0299] The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The response is then played over the speaker to the visitor. The QR code is then displayed on the digital signage.
[0300] 5. QR code payment
[0301] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[0302] 6. Recording and Monitoring
[0303] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[0304] The above is a specific embodiment of the present invention, which allows elderly people to safely and comfortably attend to visitors and furthermore allows responses that take into consideration the feelings of visitors, thereby providing a more user-friendly system.
[0305] The processing flow will be explained below.
[0306] Step 1:
[0307] Visitor Detection
[0308] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the designated detection area, the camera captures the visitor's video and transmits the information to the system.
[0309] Step 2:
[0310] Activating the microphone
[0311] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[0312] Step 3:
[0313] Digital human display and greeting
[0314] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0315] Step 4:
[0316] Capture audio data
[0317] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[0318] Step 5:
[0319] Analysis of audio data
[0320] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, for example, determining the intent "delivery."
[0321] Step 6:
[0322] Emotion recognition
[0323] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone. It recognizes emotions from the visitor's facial expressions and tone of voice.
[0324] Server: The emotion engine uses the analysis results to understand emotions such as "happy," "angry," and "sad."
[0325] Step 7:
[0326] Generating a response
[0327] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[0328] Step 8:
[0329] Sending generated responses and audio output
[0330] Server: Generates and sends the response to the device.
[0331] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[0332] Step 9:
[0333] Displaying the QR code
[0334] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[0335] Step 10:
[0336] Payment confirmation
[0337] User: A visitor scans the QR code with their smartphone and makes a payment.
[0338] Terminal: Once the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[0339] Step 11:
[0340] Record of response status
[0341] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[0342] Step 12:
[0343] Data storage and analysis
[0344] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[0345] Step 13:
[0346] User Notification
[0347] Server: Sends visitor information (image, audio, purpose, emotion, etc.) to the user and notifies them with a message such as "XX delivery has been made."
[0348] Example 2
[0349] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0350] Conventional reception systems have difficulty accurately understanding the visitor's intentions and generating appropriate responses that take their emotions into account. Furthermore, responding without accurately understanding the visitor's intentions and emotions can result in a decline in the quality of service provided to users. Furthermore, the system lacks sufficient functionality to detect suspicious behavior and maintain safety. It is necessary to solve these problems and provide efficient and safe responses to visitors.
[0351] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0352] In this invention, the server includes a means for transmitting voice data to a cloud server, a means for analyzing the voice data in the cloud server to understand the visitor's intention, and a means for performing emotion recognition based on the understood intention. This makes it possible to accurately grasp the visitor's intention and emotion and generate an appropriate response based on that. It is also possible to detect suspicious behavior in real time to ensure safety.
[0353] "Means for detecting approaching visitors" refers to a function that uses a detection device such as a camera or sensor to detect when a visitor approaches a certain location.
[0354] The "means for capturing the visitor's voice" is a function for recording the visitor's speech using a voice input device such as a microphone.
[0355] The "means for transmitting audio data to a cloud server" is a function for transmitting captured audio data to a cloud server via the Internet.
[0356] "Means for analyzing voice data on a cloud server to understand the visitor's intent" refers to a function that uses voice recognition technology on a cloud server to analyze voice data and understand what the visitor is looking for.
[0357] "Means for recognizing emotions based on understood intentions" is a function that understands the visitor's intentions and then determines their emotions from their audio and video data.
[0358] "Means for visually and audibly outputting the generated response" refers to a function that communicates the response generated by the system to the visitor using an output device such as a display or speaker.
[0359] "Means for displaying a QR code as part of a response" refers to the ability to display a QR code on a digital display as a response to a visitor.
[0360] "Means for recording interactions with visitors" refers to the system's ability to record video and audio of interactions with visitors.
[0361] "Means for saving the recorded response status on a cloud server" is a function for saving the recorded video and audio data on a cloud server.
[0362] "Means for selecting the most appropriate response" is a function that determines the response that the system deems most appropriate based on the visitor's intentions and emotions.
[0363] "Means for analyzing and detecting suspicious behavior on a cloud server" refers to a function that analyzes data stored on a cloud server and automatically detects suspicious behavior.
[0364] The present invention is a digital human-type AI reception system that automatically responds when a visitor approaches, accurately understanding the visitor's intentions and emotions and responding optimally based on the results, thereby providing a safe and comfortable experience for the visitor. Specific embodiments of the present invention are described below.
[0365] System Configuration
[0366] The system consists of the following main components:
[0367] 1. Edge AI Camera: A device that detects visitor movements and captures video.
[0368] 2. Audio pickup microphone: A device that captures the voices of visitors.
[0369] 3. Digital Signage: A device that displays a digital human and responds visually and audibly to visitors.
[0370] 4. Generative AI model: Software that runs on a cloud server, analyzes voice data, understands the visitor's intent, and generates appropriate responses.
[0371] 5. Emotion Engine: Software that runs on a cloud server and analyzes audio and video data to recognize visitors' emotions.
[0372] 6. Speaker: A device that transmits the generated voice response to the visitor.
[0373] 7. QR code display device: A device that displays QR codes for payment and verification.
[0374] 8. Cloud Server: A device that stores recorded data and analyzes suspicious behavior.
[0375] Explanation of program processing
[0376] The system works by capturing audio and video of visitors using an edge AI camera and microphone when they approach, sending the captured audio and video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[0377] Visitor detection and response
[0378] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0379] User: The visitor answers with the purpose of the call, such as "This is a courier delivery."
[0380] Speech recognition and intent understanding
[0381] Terminal: A microphone captures the visitor's voice and sends it to a cloud server.
[0382] Server: The generated AI model on the cloud server analyzes the voice data and understands the visitor's intent. For example, it determines whether the intent is "This is a courier delivery."
[0383] emotion recognition
[0384] Device: The video captured by the edge AI camera and the audio captured by the microphone are sent to the emotion engine on the cloud server.
[0385] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. For example, it recognizes the emotion "I'm in a hurry."
[0386] Response generation and display
[0387] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is in a hurry, the server generates a response such as, "You have a package to deliver. If you are in a hurry, please scan the QR code."
[0388] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[0389] QR code payment
[0390] User: Scans the displayed QR code using a smartphone and makes a payment.
[0391] Terminal: Once the payment is confirmed, the terminal will again announce to the visitor, "Payment has been confirmed. Please leave your luggage."
[0392] Recording and Monitoring
[0393] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[0394] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[0395] Specific examples
[0396] Example: This system is installed on the door of a home where elderly people live.
[0397] Visitor detection and response
[0398] Device: A delivery person arrives at the elderly person's home. When the Edge AI camera detects the delivery person, the system is activated. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[0399] User: The delivery person replies, "This is a courier delivery."
[0400] Speech recognition and intent understanding
[0401] Device: A microphone captures the voice and sends it to a cloud server. A generative AI model analyzes the voice and understands the intent of "delivery."
[0402] emotion recognition
[0403] Server: The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[0404] Response generation and display
[0405] Server: The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The speaker plays the message to the visitor. The digital sign displays the QR code.
[0406] QR code payment
[0407] User: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the delivery person is told again by voice, "The payment has been confirmed. Please leave your package."
[0408] Recording and Monitoring
[0409] Device: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[0410] Example prompt sentence:
[0411] An example of a prompt sentence that a visitor can enter into the system using the API is, "This is a courier delivery. It's urgent, so please respond quickly."
[0412] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to safely and comfortably attend to visitors, and furthermore, it is possible to respond in a way that takes into consideration the feelings of visitors, providing a more user-friendly system.
[0413] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0414] Step 1: Visitor detection
[0415] Device: The Edge AI camera detects visitor movement. The camera constantly monitors the surroundings, and when movement is detected, the system automatically activates. Specifically, it identifies human silhouettes and movements and extracts their coordinate data.
[0416] Input: Camera video data
[0417] Output: coordinate data of detected motion and visitor detection trigger signal
[0418] Step 2: Voice capture and digital human response
[0419] Terminal: The microphone begins to capture the visitor's voice. At the same time, a digital human appears on the digital signage and asks, "Welcome. How can I help you?" The captured voice is converted into audio data through a signal processing engine.
[0420] Input: Visitor voice, visitor detection trigger signal
[0421] Output: Converted audio data, activation of digital human representation
[0422] Step 3: Visitor response
[0423] User: The visitor answers with their request, such as "This is a courier delivery." The microphone captures their voice.
[0424] Input: Visitor's response voice
[0425] Output: Captured response audio data
[0426] Step 4: Sending audio data
[0427] Device: The device sends the captured audio data to a cloud server, where it is converted into packets and transmitted using a secure protocol.
[0428] Input: Captured response audio data
[0429] Output: Send audio data to cloud server
[0430] Step 5: Voice data analysis and intent understanding
[0431] Server: The AI model generated on the cloud server analyzes the voice data and understands the visitor's intent. The voice data is converted into text data through a natural language processing engine, and the intent is extracted from the text data. For example, the intent "delivery" is understood from the phrase "This is a courier delivery."
[0432] Input: Audio data sent to the cloud server
[0433] Output: Text data about the visitor's intent (e.g., "delivery")
[0434] Step 6: Sending Emotion Data
[0435] Device: Video captured by the edge AI camera and audio captured by the microphone are sent to the emotion engine on the cloud server. The video and audio data are integrated and sent as a single data packet.
[0436] Input: Video and audio data
[0437] Output: Sending sentiment analysis data to the cloud server
[0438] Step 7: Sentiment Data Analysis
[0439] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. It extracts facial features from video data and analyzes tone characteristics from audio data. For example, it recognizes emotions such as "I'm in a hurry."
[0440] Input: Video and audio data sent to the cloud server
[0441] Output: Data about visitor sentiment (e.g., "I'm in a hurry")
[0442] Step 8: Response Generation
[0443] Server: Generates appropriate responses based on information obtained through speech analysis and emotion recognition. The generative AI model generates the optimal answer based on the visitor's intent and emotion data. For example, it generates a response such as, "You're delivering a package. If it's urgent, please scan the QR code."
[0444] Input: Visitor intent data, sentiment data
[0445] Output: Text data of the response
[0446] Step 9: Response display and communication
[0447] Terminal: The generated response is transmitted to the visitor via a speaker. A QR code is displayed on the digital signage. The response text is converted into speech using a speech synthesis engine and output from the speaker.
[0448] Input: Text data of response content
[0449] Output: Audio and visual response (QR code display)
[0450] Step 10: Scan the QR code and pay
[0451] User: Scans the displayed QR code using a smartphone and makes the payment. Once the payment is completed through the payment app, a notification is automatically sent to the system.
[0452] Input: QR code
[0453] Output: Payment completion notification
[0454] Step 11: Payment confirmation and instructions
[0455] Terminal: Confirm that the payment has been completed and tell the visitor again by voice, "Payment has been confirmed. Please leave your luggage."
[0456] Input: Payment completion notification
[0457] Output: Voice notification of payment completion
[0458] Step 12: Data recording
[0459] Device: Video and audio recordings are made during the call and sent to a cloud server. The recorded data is encrypted and stored securely.
[0460] Input: Video and audio data during the call
[0461] Output: Send recorded data to cloud server
[0462] Step 13: Detect and notify suspicious behavior
[0463] Server: Analyzes stored data and detects suspicious behavior. It uses algorithms to identify anomalous patterns of behavior and notifies the user if any are detected.
[0464] Input: Stored video and audio data
[0465] Output: Suspicious behavior detected notification
[0466] (Application example 2)
[0467] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0468] Currently, there are many systems that automatically respond to visitors, but few can recognize the visitor's emotions and respond appropriately. Furthermore, only a limited number of systems have the functionality to record the results of the response and detect suspicious behavior. Therefore, there is a demand for a system that can respond to visitors efficiently and safely, and that generates responses that take the visitor's emotions into consideration.
[0469] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for detecting when a visitor is approaching, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotions, means for outputting the generated response by voice and visually displaying it, means for recording the interaction with the visitor, and means for saving the recorded interaction and analyzing it as needed. This enables appropriate interaction that takes the visitor's emotions into consideration, thereby achieving safe and efficient visitor interaction.
[0470] "Means for detecting approaching visitors" refers to devices or software that detect the movement of visitors and activate the system.
[0471] "Means for capturing visitor audio" means a microphone or recording device used to collect and record audio emitted by a visitor.
[0472] "Means for analyzing captured voice data to understand the visitor's intent" refers to algorithms or software that analyzes voice data using voice recognition technology and understands the visitor's intent and requests.
[0473] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on collected data.
[0474] "Means for outputting the generated response audibly and displaying it visually" refers to equipment or software for outputting the generated response audibly through a speaker and displaying it on a display such as digital signage.
[0475] "Visitor interaction recording means" means any device or software that records visitor interactions in video or audio format.
[0476] "Means for storing the recorded response status and analyzing it as needed" refers to software and hardware for storing and analyzing the recorded data in a database or cloud server.
[0477] "Means for displaying QR codes" refers to devices or software that generate QR codes that encode specific information and display them on a display.
[0478] "Means for detecting and responding to suspicious behavior" refers to algorithms or software that analyze recorded data, detect abnormalities or suspicious behavior, and take appropriate action.
[0479] The present invention is a system for automatically responding to visitors, recognizing their emotions, and responding appropriately. This system includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotion, means for outputting the generated response as voice and visually displaying it, means for recording the visitor interaction, and means for saving the recorded interaction and analyzing it as needed.
[0480] System Programming and Processing
[0481] The operation of this system is described as follows: This system uses the following hardware and software:
[0482] Hardware: Smartphone (camera, microphone), server, display, speaker
[0483] Software: OpenCV (image processing), SpeechRecognition (voice recognition), Transformers (generative AI model), TensorFlow (emotion recognition), qrcode (QR code generation), Firebase (data storage and analysis)
[0484] How to detect approaching visitors:
[0485] The system uses the smartphone camera to detect the movement of visitors. For example, when the smartphone camera recognizes a visitor, the system is activated.
[0486] To capture visitor audio:
[0487] The microphone on the smartphone is used to capture the visitor's voice. For example, if a visitor says, "I want to see new products," the microphone will collect that voice.
[0488] How to analyze captured audio data to understand visitor intent:
[0489] The collected voice data is converted into text data using the SpeechRecognition library. A generative AI model is then used to understand the visitor's intent. For example, a speech that says "I want to see new products" is understood to mean "I'm looking for product information."
[0490] Measures including generative AI models that generate responses based on the understood intent and sentiment of the visitor:
[0491] A response is generated based on the understood intent and the visitor's emotion (e.g., "excited") analyzed using an emotion recognition model using TensorFlow. Transformers is used as the generative AI model to input appropriate prompt sentences. For example, if the visitor is excited, the response generated will be, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[0492] A way to both speak and visually display the generated response:
[0493] The generated response is output as audio through the smartphone's speaker and displayed visually on the display, along with a QR code. For example, the QR code may be displayed along with a voice message saying, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[0494] How we record visitor interactions:
[0495] Visitor interactions are recorded in video and audio format, for example when a visitor scans a QR code.
[0496] A means to store recorded interactions and analyze them as needed:
[0497] The recorded data is stored in the cloud using Firebase and analyzed as needed to detect suspicious behavior. For example, if abnormal movements or behavior are detected, an administrator will be notified.
[0498] Examples and prompts
[0499] As a concrete example, consider a system placed at the entrance of a physical store. When a visitor approaches the store entrance, the system activates and analyzes the visitor's facial expressions and voice to respond. For example, along with a voice response such as "Welcome. How can I help you?", a response may be generated that instructs the visitor, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[0500] Prompt Sentence Examples
[0501] "The visitor's sentiment is 'Excited' and they say: 'I want to see new products.' Generate an appropriate response. Visitor: 'I want to see new products.' Example of an appropriate response: 'Here is a list of new products. Choose your favorite and scan the QR code.'"
[0502] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0503] Step 1:
[0504] A means (device) of detecting approaching visitors
[0505] The system uses the smartphone camera to detect the movement of visitors. Specifically, it uses the OpenCV library to analyze the camera footage and perform motion detection. When motion is detected, the system is activated.
[0506] Input: Camera video data
[0507] Output: A signal that a visitor is approaching
[0508] Step 2:
[0509] A means (device) for capturing visitor audio
[0510] It uses the device's microphone to capture the visitor's voice, specifically by using the SpeechRecognition library to collect the voice data.
[0511] Input: Visitor utterance
[0512] Output: Audio data
[0513] Step 3:
[0514] A means of analyzing voice data and understanding the visitor's intent (server)
[0515] The captured voice data is sent to the server, where it is converted into text using the SpeechRecognition library, and then generative AI models (Transformers) are used to understand the visitor's intent.
[0516] Input: Audio data
[0517] Output: Text data containing visitor intent
[0518] Step 4:
[0519] A means of analyzing visitor sentiment (server)
[0520] Using video data captured on the server, we analyze visitors' emotions using a TensorFlow model, specifically identifying emotions from facial expressions and tone of voice.
[0521] Input: Video and audio data
[0522] Output: Visitor sentiment data
[0523] Step 5:
[0524] A means (server) to generate responses based on the understood intent and emotions
[0525] Using generative AI models (Transformers), it generates appropriate responses based on the visitor's intent and emotions. For example, if the visitor is excited, it might generate a response like, "Here's a list of new products. Choose your favorite and scan the QR code."
[0526] Input: Text data containing intent, emotion data
[0527] Output: The generated response text
[0528] Step 6:
[0529] A means (terminal) to output the generated response as voice and display it visually
[0530] The generated response text is synthesized into speech and output as voice through the smartphone speaker. The response text including a QR code is also displayed on the display. The QR code is generated using the qrcode library.
[0531] Input: Generated response text
[0532] Output: Audio output, visual display with QR code
[0533] Step 7:
[0534] A means (terminal) for recording visitor interactions
[0535] Record visitor interactions in video and audio format using the smartphone camera and microphone.
[0536] Input: Video and audio when greeting a visitor
[0537] Output: Recorded video and audio data
[0538] Step 8:
[0539] A means (server) to store the recorded response status and analyze it as needed
[0540] The recorded video and audio data is stored on the Firebase cloud server. If necessary, the stored data is analyzed, and if any suspicious activity is detected, an administrator is notified.
[0541] Input: Recorded video and audio data
[0542] Output: Data stored in the cloud, suspicious behavior detection results
[0543] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0544] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0545] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0546] [Second embodiment]
[0547] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0548] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0549] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0550] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0551] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0552] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0553] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0554] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0555] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0556] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0557] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0558] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0559] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to respond to visitors efficiently and safely. The following describes specific embodiments of the present invention.
[0560] System Configuration
[0561] The system consists of the following main components:
[0562] 1. Edge AI Camera: Detects visitor movements and captures footage.
[0563] 2. Audio pickup microphone: Captures the voice of visitors.
[0564] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[0565] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0566] 5. Speaker: Communicates the generated voice response to the visitor.
[0567] 6. QR code display device: Displays a QR code for payment and verification.
[0568] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[0569] Explanation of program processing
[0570] The system works by using an edge AI camera and a microphone to capture the visitor's voice and video when they approach and sending it to a cloud server.
[0571] Visitor detection and response
[0572] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0573] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[0574] Speech recognition and intent understanding
[0575] Device: A microphone captures the visitor's voice and sends it to the generative AI model.
[0576] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, "delivery."
[0577] Response generation and display
[0578] Server: Generates an appropriate response based on the information obtained from speech analysis, for example, "You're delivering a package. Please scan the QR code."
[0579] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[0580] QR code payment
[0581] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[0582] Terminal: Confirm that the payment has been completed and repeat the voice message, "Payment has been confirmed. Please place your luggage."
[0583] Recording and Monitoring
[0584] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[0585] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[0586] Specific examples
[0587] Examples:
[0588] This system will be installed on the doors of homes where elderly people live.
[0589] 1. Visitor detection and response
[0590] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[0591] 2. Speech Recognition and Intent Understanding
[0592] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[0593] 3. Response Generation and Display
[0594] The generative AI model generates a response saying, "This is a package delivery. Please scan the QR code," and the response is relayed to the visitor through a speaker. The QR code is then displayed on a digital sign.
[0595] 4. QR code payment
[0596] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[0597] 5. Recording and Monitoring
[0598] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[0599] The above is a specific embodiment for carrying out the present invention, which allows elderly people to safely and comfortably receive visitors and live with peace of mind even when they are away or busy.
[0600] The processing flow will be explained below.
[0601] Step 1:
[0602] Visitor Detection
[0603] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the detection area, the camera captures the visitor's video and transmits the information to the system.
[0604] Step 2:
[0605] Activating the microphone
[0606] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[0607] Step 3:
[0608] Digital human display and greeting
[0609] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0610] Step 4:
[0611] Capture audio data
[0612] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[0613] Step 5:
[0614] Analysis of audio data
[0615] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, such as "delivery."
[0616] Step 6:
[0617] Generating a response
[0618] Server: Generates an appropriate response based on the visitor's intent. For example, "You're about to deliver a package. Please scan the QR code."
[0619] Step 7:
[0620] Sending generated responses and audio output
[0621] Server: Generates and sends the response to the device.
[0622] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[0623] Step 8:
[0624] Displaying the QR code
[0625] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[0626] Step 9:
[0627] Payment confirmation
[0628] User: A visitor scans the QR code with their smartphone and makes a payment.
[0629] Terminal: If the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[0630] Step 10:
[0631] Record of response status
[0632] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[0633] Step 11:
[0634] Data storage and analysis
[0635] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[0636] Step 12:
[0637] User Notification
[0638] Server: Sends visitor information (images, audio, purpose, etc.) to the user and notifies them with a message such as "XX Takkyubin has been delivered."
[0639] Example 1
[0640] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0641] In modern society, responding to visitors needs to be done efficiently and safely. In particular, for elderly people and those living alone, face-to-face interactions with visitors are often a burden. Furthermore, when the home is busy or out of the home, prompt and accurate responses are required, and it is also important to detect suspicious behavior. To solve these issues, a system is needed that can automatically detect visitor movements, understand the visitor's intentions, and provide an appropriate response.
[0642] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0643] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intention, means for outputting the generated response as voice, means for displaying visual information to the visitor, means for recording the visitor's interaction, and means for saving the recorded interaction. This not only enables efficient interaction with visitors, but also enables the visitor's intention to be understood quickly and accurately, enabling appropriate interaction even when the visitor is absent or busy. Furthermore, by detecting suspicious behavior, safety is improved.
[0644] "Means for detecting approaching visitors" refers to devices such as sensors and cameras that detect visitor movement, as well as software that controls them.
[0645] "Means for capturing visitor audio" refers to a microphone for recording the audio made by the visitor, and any equipment or software for appropriately processing that audio data.
[0646] "Means for analyzing voice data to understand visitor intent" refers to algorithms and software for analyzing captured voice data and understanding its content, primarily using generative AI models.
[0647] "Means for generating responses based on understood intent" refers to software and algorithms that understand the visitor's intent and generate appropriate responses based on that content.
[0648] The "means for outputting the generated response as voice" refers to a speaker for outputting the generated response as voice, and a device and software for controlling the speaker.
[0649] "Means for displaying visual information to visitors" refers to devices such as displays and digital signage that provide visual information to visitors, as well as software for controlling them.
[0650] "Means for recording interactions with visitors" refers to devices such as cameras and microphones for recording and recording interactions with visitors and the situation, as well as software for controlling them.
[0651] "Means for storing recorded responses" refers to a storage device for safely storing video and audio data, and software for managing that data.
[0652] System configuration and operation overview
[0653] This system is a digital human-type AI reception system for automatically responding to visitors. The system consists of the following main hardware and software components:
[0654] 1. Edge AI Camera: Detects visitor movements and captures footage.
[0655] 2. Audio pickup microphone: Captures the voice of visitors.
[0656] 3. Digital Signage: Displaying a digital human and responding visually and audibly to visitors.
[0657] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0658] 5. Speaker: Communicates the generated voice response to the visitor.
[0659] 6. QR code display device: Displays a QR code for payment and verification.
[0660] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[0661] System operation details
[0662] Visitor Detection
[0663] Terminal: When the Edge AI camera detects a visitor's movement, the system automatically wakes up. At the same time, the audio pickup microphone activates and captures the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0664] Audio capture and transmission
[0665] Terminal: The visitor's voice (e.g., "This is a parcel delivery") is captured by a microphone and sent to a cloud server. Since the voice data is processed in real time, low-latency data transfer is required.
[0666] Voice analysis and intent understanding
[0667] Server: The cloud server analyzes the received voice data using a generative AI model to understand the visitor's intent. For example, it can extract the intent "delivery" from the voice data "This is a courier delivery."
[0668] Response generation and display
[0669] Server: The generative AI model generates an appropriate response based on the analysis results, for example, "This is a package delivery. Please scan the QR code."
[0670] Terminal: The text response sent from the server is communicated to the visitor using a digital sign and speaker. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the QR code is displayed on the digital sign.
[0671] QR code payment
[0672] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[0673] Terminal: Once the payment is confirmed, the visitor will be notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[0674] Recording and Monitoring
[0675] Terminal: Records and records interactions with visitors and sends the data to a cloud server.
[0676] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[0677] Example operation
[0678] Consider the case where this system is installed on the door of a home where elderly people live.
[0679] 1. Visitor Detection: When a courier approaches the door, the Edge AI camera detects the movement and the system is activated. A digital human appears and asks the courier, "Welcome. How can I help you?"
[0680] 2. Voice capture and transmission: The delivery person responds, "This is a courier delivery," and the voice is captured by the microphone and transmitted to the cloud server.
[0681] 3. Speech analysis and intent understanding: The generative AI model analyzes the voice data and understands the intent of "delivery."
[0682] 4. Response generation and display: The generative AI model generates a response such as "This is a package delivery. Please scan the QR code," plays the audio over the speaker, and displays the QR code on the digital signage.
[0683] 5. QR code payment: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, they will receive a voice message saying, "Payment has been confirmed. Please leave your package."
[0684] 6. Recording and monitoring: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[0685] As a result, this system can respond to visitors efficiently and safely, and can provide safe and secure living support for elderly people and those living alone in their homes.
[0686] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0687] Step 1: Visitor detection
[0688] Terminal: When the Edge AI camera detects a visitor's movement, the system wakes up. The audio pickup microphone activates and prepares to capture the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can we help you?"
[0689] Input: Visitor motion detection (Edge AI camera)
[0690] Output: System startup, digital human display, greeting message output
[0691] Step 2: Capture audio
[0692] Terminal: A microphone captures the visitor's voice, such as "This is a courier delivery," and the voice data is converted into a digital format.
[0693] Input: Visitor's voice
[0694] Output: Digital audio data
[0695] Step 3: Audio transmission and analysis
[0696] Terminal: The captured audio data is sent to the cloud server in real time.
[0697] Server: Analyzes the received voice data using a generative AI model to understand the visitor's intent. Converts the voice data into text and uses natural language processing technology to extract the intent. For example, understand the intent of "delivery" from the voice saying "This is a courier delivery."
[0698] Input: Digital audio data
[0699] Output: Text data containing intent
[0700] Step 4: Response Generation
[0701] Server: The generative AI model generates an appropriate response based on the analysis results, for example, a text response such as "This is a package delivery. Please scan the QR code."
[0702] Input: Text data containing visitor intent
[0703] Output: Response text
[0704] Step 5: Response display and audio output
[0705] Terminal: The response text received from the server is communicated to the visitor via a speaker and digital signage. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the digital signage displays the QR code.
[0706] Input: Response text
[0707] Output: Voice response, QR code display
[0708] Step 6: QR code payment
[0709] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[0710] Terminal: The QR code is scanned to confirm the payment has been completed. Once the payment is complete, the visitor is notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[0711] Input: QR code scan, payment information
[0712] Output: Payment completion notification
[0713] Step 7: Record and monitor
[0714] Terminal: Records and records interactions with visitors and sends the data, including video and audio data, to a cloud server.
[0715] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[0716] Input: Video data, audio data
[0717] Output: Saved data, suspicious behavior notification
[0718] (Application example 1)
[0719] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0720] Modern stores are required to respond to visitors quickly and accurately, but this places a heavy burden on employees, making it difficult for visitors to receive satisfactory service. It is also important to detect suspicious behavior early and respond appropriately. The present invention aims to solve these problems and respond to visitors efficiently and safely. Another objective of the present invention is to enable employees to smoothly confirm visitor responses and improve the efficiency of store operations.
[0721] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0722] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intent, means for generating a response based on the understood intent, means for outputting the generated response as voice, means for recording the interaction with the visitor, means for saving the recorded interaction, and means for an employee wearing smart glasses to visually confirm the response. This allows employees to instantly understand the visitor's intent through the smart glasses and respond promptly and appropriately. Furthermore, recording and saving the interaction makes it easy to review later and detect suspicious behavior.
[0723] A "means for detecting approaching visitors" is a device or method that senses the movement of a visitor and triggers activation of the system.
[0724] A "visitor voice capturing means" is a device or method that collects voice from a visitor and processes it as digital data.
[0725] The "means for analyzing voice data and understanding the visitor's intent" refers to a device or method for analyzing collected voice data and understanding the visitor's requests and questions.
[0726] The "means for generating a response based on the understood intent" is a device or method that automatically generates an appropriate response based on the analysis results.
[0727] The "means for outputting the generated response by voice" is a device or method for transmitting the generated response to the visitor as voice.
[0728] "Means for recording visitor interactions" refers to a device or method for saving the interaction and interaction with visitors as digital data.
[0729] "Means for storing recorded response situations" refers to a device or method for durably storing recorded digital data.
[0730] A "means for visually confirming a response by an employee equipped with smart glasses" is a device or method for visually confirming a response generated through smart glasses worn by an employee.
[0731] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and an embodiment thereof is shown based on an application example using smart glasses worn by employees.
[0732] System configuration:
[0733] The system consists of the following main components:
[0734] 1. Edge AI Camera: When a visitor enters the store, it detects their movement and activates the entire system.
[0735] 2. Audio pickup microphone: Collects visitors' voices and processes them as digital data.
[0736] 3. Smart glasses: Devices worn by employees to visually and audibly confirm the visitor's intentions.
[0737] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0738] 5. Speaker: Outputs the generated response as audio and conveys it to the visitor.
[0739] 6. Cloud server: Stores recorded responses and detects suspicious behavior.
[0740] Explanation of program operation:
[0741] The server controls the edge AI camera, microphone, smart glasses, generative AI model, speaker, and cloud server to enable interaction with visitors. The main processing flow is as follows:
[0742] Visitor Detection:
[0743] When the Edge AI camera detects a visitor's movement, the system is activated, the microphone captures the visitor's voice, and a digital human appears in the smart glasses and asks the visitor, "Welcome. How can I help you?"
[0744] Speech Recognition and Intent Understanding:
[0745] The system captures audio data with a microphone and sends it to a generative AI model, which then analyzes the audio data to understand the visitor's intent.
[0746] For example, if a visitor says, "Please tell me where the new products are," the generative AI model understands the intent "new products" and generates the response, "The new products corner is in the back right."
[0747] Response generation and display:
[0748] The generative AI model generates an appropriate response, which is output through the speaker and also displayed on the smart glasses' display.
[0749] Record of response:
[0750] Visitor interactions are videotaped and stored on a cloud server, and if any suspicious activity is detected, a notification is sent to the administrator.
[0751] Hardware and software used:
[0752] Hardware: Edge AI camera (general security camera), sound pickup microphone (standalone microphone), smart glasses (e.g., smart glasses), speaker.
[0753] Software: speech recognition models (e.g., Google Speech-to-Text API), generative AI models (e.g., OpenAI GPT), and cloud storage (e.g., Google Cloud).
[0754] Example prompt sentence:
[0755] Input speech data: "Please tell me where the new products are located."
[0756] Prompt: "Based on the audio data, analyze the visitor's intent and generate an appropriate response. For example, a response to a question about the location of a new product."
[0757] Thus, the present invention utilizes cutting edge technology such as smart glasses to provide a system that can efficiently and safely serve visitors.
[0758] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0759] Step 1:
[0760] Visitor Detection
[0761] The device's edge AI camera detects the visitor's movements. The motion detection activates the system, which then captures the visitor's video data. The input is the visitor's movements, and the output is the video data and a signal to activate the system.
[0762] Step 2:
[0763] Audio Capture
[0764] The device's microphone captures the visitor's voice, and the collected voice data is sent to the system. The input is the visitor's voice, and the output is digital voice data.
[0765] Step 3:
[0766] Analysis of audio data
[0767] The server's generated AI model analyzes the collected voice data and understands the visitor's intent. The input is digital voice data, and the output is the visitor's intent as an analysis result. A voice recognition model is used for the analysis.
[0768] Step 4:
[0769] Response Generation
[0770] The server's generative AI model generates an appropriate response based on the visitor's intent. The input is the visitor's intent as a result of analysis, and the output is the generated response text. For example, if the question is about the location of new products, the generated response will be "The new products corner is in the back right."
[0771] Step 5:
[0772] Display and speak responses
[0773] The smart glasses on the terminal visually display the generated response, and the speaker outputs the response aloud. The input is the generated response text, and the output is the visual display and audio output. Employees can respond to visitors' questions instantly through the smart glasses.
[0774] Step 6:
[0775] Record of response status
[0776] The device records the conversation with the visitor and sends it to the cloud server. The input is the audio and video data of the conversation with the visitor, and the output is the recorded data sent to the cloud server.
[0777] Step 7:
[0778] Data storage and analysis
[0779] The cloud server stores the recorded response data and analyzes suspicious behavior. The input is video and audio data, and the output is the analysis result, detecting suspicious behavior and notifying the administrator.
[0780] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0781] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to efficiently and safely serve visitors while recognizing their emotions and responding appropriately. A specific embodiment of the present invention will be described below.
[0782] System Configuration
[0783] The system consists of the following main components:
[0784] 1. Edge AI Camera: Detects visitor movements and captures footage.
[0785] 2. Audio pickup microphone: Captures the voice of visitors.
[0786] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[0787] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[0788] 5. Emotion Engine: Analyzes audio and video data to recognize visitors' emotions.
[0789] 6. Speaker: Communicates the generated voice response to the visitor.
[0790] 7. QR code display device: Displays a QR code for payment and verification.
[0791] 8. Cloud server: Stores recorded data and analyzes suspicious behavior.
[0792] Explanation of program processing
[0793] The system works by capturing the visitor's voice and video using an edge AI camera and microphone when the visitor approaches, sending the captured video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[0794] Visitor detection and response
[0795] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0796] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[0797] Speech recognition and intent understanding
[0798] Device: A microphone captures the visitor's voice and sends it to a cloud server running a generative AI model.
[0799] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, determining the intent to "deliver."
[0800] emotion recognition
[0801] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone.
[0802] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice, such as "happy," "angry," or "sad."
[0803] Response generation and display
[0804] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[0805] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[0806] QR code payment
[0807] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[0808] Terminal: Confirms that the payment has been completed and repeats the message "Payment has been confirmed. Please place your luggage."
[0809] Recording and Monitoring
[0810] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[0811] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[0812] Specific examples
[0813] Examples:
[0814] This system will be installed on the doors of homes where elderly people live.
[0815] 1. Visitor detection and response
[0816] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[0817] 2. Speech Recognition and Intent Understanding
[0818] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[0819] 3. Emotion recognition
[0820] The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[0821] 4. Response Generation and Display
[0822] The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The response is then played over the speaker to the visitor. The QR code is then displayed on the digital signage.
[0823] 5. QR code payment
[0824] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[0825] 6. Recording and Monitoring
[0826] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[0827] The above is a specific embodiment of the present invention, which allows elderly people to safely and comfortably attend to visitors and furthermore allows responses that take into consideration the feelings of visitors, thereby providing a more user-friendly system.
[0828] The processing flow will be explained below.
[0829] Step 1:
[0830] Visitor Detection
[0831] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the designated detection area, the camera captures the visitor's video and transmits the information to the system.
[0832] Step 2:
[0833] Activating the microphone
[0834] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[0835] Step 3:
[0836] Digital human display and greeting
[0837] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0838] Step 4:
[0839] Capture audio data
[0840] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[0841] Step 5:
[0842] Analysis of audio data
[0843] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, for example, determining the intent "delivery."
[0844] Step 6:
[0845] Emotion recognition
[0846] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone. It recognizes emotions from the visitor's facial expressions and tone of voice.
[0847] Server: The emotion engine uses the analysis results to understand emotions such as "happy," "angry," and "sad."
[0848] Step 7:
[0849] Generating a response
[0850] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[0851] Step 8:
[0852] Sending generated responses and audio output
[0853] Server: Generates and sends the response to the device.
[0854] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[0855] Step 9:
[0856] Displaying the QR code
[0857] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[0858] Step 10:
[0859] Payment confirmation
[0860] User: A visitor scans the QR code with their smartphone and makes a payment.
[0861] Terminal: Once the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[0862] Step 11:
[0863] Record of response status
[0864] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[0865] Step 12:
[0866] Data storage and analysis
[0867] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[0868] Step 13:
[0869] User Notification
[0870] Server: Sends visitor information (image, audio, purpose, emotion, etc.) to the user and notifies them with a message such as "XX delivery has been made."
[0871] Example 2
[0872] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0873] Conventional reception systems have difficulty accurately understanding the visitor's intentions and generating appropriate responses that take their emotions into account. Furthermore, responding without accurately understanding the visitor's intentions and emotions can result in a decline in the quality of service provided to users. Furthermore, the system lacks sufficient functionality to detect suspicious behavior and maintain safety. It is necessary to solve these problems and provide efficient and safe responses to visitors.
[0874] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0875] In this invention, the server includes a means for transmitting voice data to a cloud server, a means for analyzing the voice data in the cloud server to understand the visitor's intention, and a means for performing emotion recognition based on the understood intention. This makes it possible to accurately grasp the visitor's intention and emotion and generate an appropriate response based on that. It is also possible to detect suspicious behavior in real time to ensure safety.
[0876] "Means for detecting approaching visitors" refers to a function that uses a detection device such as a camera or sensor to detect when a visitor approaches a certain location.
[0877] The "means for capturing the visitor's voice" is a function for recording the visitor's speech using a voice input device such as a microphone.
[0878] The "means for transmitting audio data to a cloud server" is a function for transmitting captured audio data to a cloud server via the Internet.
[0879] "Means for analyzing voice data on a cloud server to understand the visitor's intent" refers to a function that uses voice recognition technology on a cloud server to analyze voice data and understand what the visitor is looking for.
[0880] "Means for recognizing emotions based on understood intentions" is a function that understands the visitor's intentions and then determines their emotions from their audio and video data.
[0881] "Means for visually and audibly outputting the generated response" refers to a function that communicates the response generated by the system to the visitor using an output device such as a display or speaker.
[0882] "Means for displaying a QR code as part of a response" refers to the ability to display a QR code on a digital display as a response to a visitor.
[0883] "Means for recording interactions with visitors" refers to the system's ability to record video and audio of interactions with visitors.
[0884] "Means for saving the recorded response status on a cloud server" is a function for saving the recorded video and audio data on a cloud server.
[0885] "Means for selecting the most appropriate response" is a function that determines the response that the system deems most appropriate based on the visitor's intentions and emotions.
[0886] "Means for analyzing and detecting suspicious behavior on a cloud server" refers to a function that analyzes data stored on a cloud server and automatically detects suspicious behavior.
[0887] The present invention is a digital human-type AI reception system that automatically responds when a visitor approaches, accurately understanding the visitor's intentions and emotions and responding optimally based on the results, thereby providing a safe and comfortable experience for the visitor. Specific embodiments of the present invention are described below.
[0888] System Configuration
[0889] The system consists of the following main components:
[0890] 1. Edge AI Camera: A device that detects visitor movements and captures video.
[0891] 2. Audio pickup microphone: A device that captures the voices of visitors.
[0892] 3. Digital Signage: A device that displays a digital human and responds visually and audibly to visitors.
[0893] 4. Generative AI model: Software that runs on a cloud server, analyzes voice data, understands the visitor's intent, and generates appropriate responses.
[0894] 5. Emotion Engine: Software that runs on a cloud server and analyzes audio and video data to recognize visitors' emotions.
[0895] 6. Speaker: A device that transmits the generated voice response to the visitor.
[0896] 7. QR code display device: A device that displays QR codes for payment and verification.
[0897] 8. Cloud Server: A device that stores recorded data and analyzes suspicious behavior.
[0898] Explanation of program processing
[0899] The system works by capturing audio and video of visitors using an edge AI camera and microphone when they approach, sending the captured audio and video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[0900] Visitor detection and response
[0901] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[0902] User: The visitor answers with the purpose of the call, such as "This is a courier delivery."
[0903] Speech recognition and intent understanding
[0904] Terminal: A microphone captures the visitor's voice and sends it to a cloud server.
[0905] Server: The generated AI model on the cloud server analyzes the voice data and understands the visitor's intent. For example, it determines whether the intent is "This is a courier delivery."
[0906] emotion recognition
[0907] Device: The video captured by the edge AI camera and the audio captured by the microphone are sent to the emotion engine on the cloud server.
[0908] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. For example, it recognizes the emotion "I'm in a hurry."
[0909] Response generation and display
[0910] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is in a hurry, the server generates a response such as, "You have a package to deliver. If you are in a hurry, please scan the QR code."
[0911] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[0912] QR code payment
[0913] User: Scans the displayed QR code using a smartphone and makes a payment.
[0914] Terminal: Once the payment is confirmed, the terminal will again announce to the visitor, "Payment has been confirmed. Please leave your luggage."
[0915] Recording and Monitoring
[0916] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[0917] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[0918] Specific examples
[0919] Example: This system is installed on the door of a home where elderly people live.
[0920] Visitor detection and response
[0921] Device: A delivery person arrives at the elderly person's home. When the Edge AI camera detects the delivery person, the system is activated. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[0922] User: The delivery person replies, "This is a courier delivery."
[0923] Speech recognition and intent understanding
[0924] Device: A microphone captures the voice and sends it to a cloud server. A generative AI model analyzes the voice and understands the intent of "delivery."
[0925] emotion recognition
[0926] Server: The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[0927] Response generation and display
[0928] Server: The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The speaker plays the message to the visitor. The digital sign displays the QR code.
[0929] QR code payment
[0930] User: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the delivery person is told again by voice, "The payment has been confirmed. Please leave your package."
[0931] Recording and Monitoring
[0932] Device: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[0933] Example prompt sentence:
[0934] An example of a prompt sentence that a visitor can enter into the system using the API is, "This is a courier delivery. It's urgent, so please respond quickly."
[0935] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to safely and comfortably attend to visitors, and furthermore, it is possible to respond in a way that takes into consideration the feelings of visitors, providing a more user-friendly system.
[0936] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0937] Step 1: Visitor detection
[0938] Device: The Edge AI camera detects visitor movement. The camera constantly monitors the surroundings, and when movement is detected, the system automatically activates. Specifically, it identifies human silhouettes and movements and extracts their coordinate data.
[0939] Input: Camera video data
[0940] Output: coordinate data of detected motion and visitor detection trigger signal
[0941] Step 2: Voice capture and digital human response
[0942] Terminal: The microphone begins to capture the visitor's voice. At the same time, a digital human appears on the digital signage and asks, "Welcome. How can I help you?" The captured voice is converted into audio data through a signal processing engine.
[0943] Input: Visitor voice, visitor detection trigger signal
[0944] Output: Converted audio data, activation of digital human representation
[0945] Step 3: Visitor response
[0946] User: The visitor answers with their request, such as "This is a courier delivery." The microphone captures their voice.
[0947] Input: Visitor's response voice
[0948] Output: Captured response audio data
[0949] Step 4: Sending audio data
[0950] Device: The device sends the captured audio data to a cloud server, where it is converted into packets and transmitted using a secure protocol.
[0951] Input: Captured response audio data
[0952] Output: Send audio data to cloud server
[0953] Step 5: Voice data analysis and intent understanding
[0954] Server: The AI model generated on the cloud server analyzes the voice data and understands the visitor's intent. The voice data is converted into text data through a natural language processing engine, and the intent is extracted from the text data. For example, the intent "delivery" is understood from the phrase "This is a courier delivery."
[0955] Input: Audio data sent to the cloud server
[0956] Output: Text data about the visitor's intent (e.g., "delivery")
[0957] Step 6: Sending Emotion Data
[0958] Device: Video captured by the edge AI camera and audio captured by the microphone are sent to the emotion engine on the cloud server. The video and audio data are integrated and sent as a single data packet.
[0959] Input: Video and audio data
[0960] Output: Sending sentiment analysis data to the cloud server
[0961] Step 7: Sentiment Data Analysis
[0962] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. It extracts facial features from video data and analyzes tone characteristics from audio data. For example, it recognizes emotions such as "I'm in a hurry."
[0963] Input: Video and audio data sent to the cloud server
[0964] Output: Data about visitor sentiment (e.g., "I'm in a hurry")
[0965] Step 8: Response Generation
[0966] Server: Generates appropriate responses based on information obtained through speech analysis and emotion recognition. The generative AI model generates the optimal answer based on the visitor's intent and emotion data. For example, it generates a response such as, "You're delivering a package. If it's urgent, please scan the QR code."
[0967] Input: Visitor intent data, sentiment data
[0968] Output: Text data of the response
[0969] Step 9: Response display and communication
[0970] Terminal: The generated response is transmitted to the visitor via a speaker. A QR code is displayed on the digital signage. The response text is converted into speech using a speech synthesis engine and output from the speaker.
[0971] Input: Text data of response content
[0972] Output: Audio and visual response (QR code display)
[0973] Step 10: Scan the QR code and pay
[0974] User: Scans the displayed QR code using a smartphone and makes the payment. Once the payment is completed through the payment app, a notification is automatically sent to the system.
[0975] Input: QR code
[0976] Output: Payment completion notification
[0977] Step 11: Payment confirmation and instructions
[0978] Terminal: Confirm that the payment has been completed and tell the visitor again by voice, "Payment has been confirmed. Please leave your luggage."
[0979] Input: Payment completion notification
[0980] Output: Voice notification of payment completion
[0981] Step 12: Data recording
[0982] Device: Video and audio recordings are made during the call and sent to a cloud server. The recorded data is encrypted and stored securely.
[0983] Input: Video and audio data during the call
[0984] Output: Send recorded data to cloud server
[0985] Step 13: Detect and notify suspicious behavior
[0986] Server: Analyzes stored data and detects suspicious behavior. It uses algorithms to identify anomalous patterns of behavior and notifies the user if any are detected.
[0987] Input: Stored video and audio data
[0988] Output: Suspicious behavior detected notification
[0989] (Application example 2)
[0990] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0991] Currently, there are many systems that automatically respond to visitors, but few can recognize the visitor's emotions and respond appropriately. Furthermore, only a limited number of systems have the functionality to record the results of the response and detect suspicious behavior. Therefore, there is a demand for a system that can respond to visitors efficiently and safely, and that generates responses that take the visitor's emotions into consideration.
[0992] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for detecting when a visitor is approaching, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotions, means for outputting the generated response by voice and visually displaying it, means for recording the interaction with the visitor, and means for saving the recorded interaction and analyzing it as needed. This enables appropriate interaction that takes the visitor's emotions into consideration, thereby achieving safe and efficient visitor interaction.
[0993] "Means for detecting approaching visitors" refers to devices or software that detect the movement of visitors and activate the system.
[0994] "Means for capturing visitor audio" means a microphone or recording device used to collect and record audio emitted by a visitor.
[0995] "Means for analyzing captured voice data to understand the visitor's intent" refers to algorithms or software that analyzes voice data using voice recognition technology and understands the visitor's intent and requests.
[0996] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on collected data.
[0997] "Means for outputting the generated response audibly and displaying it visually" refers to equipment or software for outputting the generated response audibly through a speaker and displaying it on a display such as digital signage.
[0998] "Visitor interaction recording means" means any device or software that records visitor interactions in video or audio format.
[0999] "Means for storing the recorded response status and analyzing it as needed" refers to software and hardware for storing and analyzing the recorded data in a database or cloud server.
[1000] "Means for displaying QR codes" refers to devices or software that generate QR codes that encode specific information and display them on a display.
[1001] "Means for detecting and responding to suspicious behavior" refers to algorithms or software that analyze recorded data, detect abnormalities or suspicious behavior, and take appropriate action.
[1002] The present invention is a system for automatically responding to visitors, recognizing their emotions, and responding appropriately. This system includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotion, means for outputting the generated response as voice and visually displaying it, means for recording the visitor interaction, and means for saving the recorded interaction and analyzing it as needed.
[1003] System Programming and Processing
[1004] The operation of this system is described as follows: This system uses the following hardware and software:
[1005] Hardware: Smartphone (camera, microphone), server, display, speaker
[1006] Software: OpenCV (image processing), SpeechRecognition (voice recognition), Transformers (generative AI model), TensorFlow (emotion recognition), qrcode (QR code generation), Firebase (data storage and analysis)
[1007] How to detect approaching visitors:
[1008] The system uses the smartphone camera to detect the movement of visitors. For example, when the smartphone camera recognizes a visitor, the system is activated.
[1009] To capture visitor audio:
[1010] The microphone on the smartphone is used to capture the visitor's voice. For example, if a visitor says, "I want to see new products," the microphone will collect that voice.
[1011] How to analyze captured audio data to understand visitor intent:
[1012] The collected voice data is converted into text data using the SpeechRecognition library. A generative AI model is then used to understand the visitor's intent. For example, a speech that says "I want to see new products" is understood to mean "I'm looking for product information."
[1013] Measures including generative AI models that generate responses based on the understood intent and sentiment of the visitor:
[1014] A response is generated based on the understood intent and the visitor's emotion (e.g., "excited") analyzed using an emotion recognition model using TensorFlow. Transformers is used as the generative AI model to input appropriate prompt sentences. For example, if the visitor is excited, the response generated will be, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[1015] A way to both speak and visually display the generated response:
[1016] The generated response is output as audio through the smartphone's speaker and displayed visually on the display, along with a QR code. For example, the QR code may be displayed along with a voice message saying, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[1017] How we record visitor interactions:
[1018] Visitor interactions are recorded in video and audio format, for example when a visitor scans a QR code.
[1019] A means to store recorded interactions and analyze them as needed:
[1020] The recorded data is stored in the cloud using Firebase and analyzed as needed to detect suspicious behavior. For example, if abnormal movements or behavior are detected, an administrator will be notified.
[1021] Examples and prompts
[1022] As a concrete example, consider a system placed at the entrance of a physical store. When a visitor approaches the store entrance, the system activates and analyzes the visitor's facial expressions and voice to respond. For example, along with a voice response such as "Welcome. How can I help you?", a response may be generated that instructs the visitor, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[1023] Prompt Sentence Examples
[1024] "The visitor's sentiment is 'Excited' and they say: 'I want to see new products.' Generate an appropriate response. Visitor: 'I want to see new products.' Example of an appropriate response: 'Here is a list of new products. Choose your favorite and scan the QR code.'"
[1025] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1026] Step 1:
[1027] A means (device) of detecting approaching visitors
[1028] The system uses the smartphone camera to detect the movement of visitors. Specifically, it uses the OpenCV library to analyze the camera footage and perform motion detection. When motion is detected, the system is activated.
[1029] Input: Camera video data
[1030] Output: A signal that a visitor is approaching
[1031] Step 2:
[1032] A means (device) for capturing visitor audio
[1033] It uses the device's microphone to capture the visitor's voice, specifically by using the SpeechRecognition library to collect the voice data.
[1034] Input: Visitor utterance
[1035] Output: Audio data
[1036] Step 3:
[1037] A means of analyzing voice data and understanding the visitor's intent (server)
[1038] The captured voice data is sent to the server, where it is converted into text using the SpeechRecognition library, and then generative AI models (Transformers) are used to understand the visitor's intent.
[1039] Input: Audio data
[1040] Output: Text data containing visitor intent
[1041] Step 4:
[1042] A means of analyzing visitor sentiment (server)
[1043] Using video data captured on the server, we analyze visitors' emotions using a TensorFlow model, specifically identifying emotions from facial expressions and tone of voice.
[1044] Input: Video and audio data
[1045] Output: Visitor sentiment data
[1046] Step 5:
[1047] A means (server) to generate responses based on the understood intent and emotions
[1048] Using generative AI models (Transformers), it generates appropriate responses based on the visitor's intent and emotions. For example, if the visitor is excited, it might generate a response like, "Here's a list of new products. Choose your favorite and scan the QR code."
[1049] Input: Text data containing intent, emotion data
[1050] Output: The generated response text
[1051] Step 6:
[1052] A means (terminal) to output the generated response as voice and display it visually
[1053] The generated response text is synthesized into speech and output as voice through the smartphone speaker. The response text including a QR code is also displayed on the display. The QR code is generated using the qrcode library.
[1054] Input: Generated response text
[1055] Output: Audio output, visual display with QR code
[1056] Step 7:
[1057] A means (terminal) for recording visitor interactions
[1058] Record visitor interactions in video and audio format using the smartphone camera and microphone.
[1059] Input: Video and audio when greeting a visitor
[1060] Output: Recorded video and audio data
[1061] Step 8:
[1062] A means (server) to store the recorded response status and analyze it as needed
[1063] The recorded video and audio data is stored on the Firebase cloud server. If necessary, the stored data is analyzed, and if any suspicious activity is detected, an administrator is notified.
[1064] Input: Recorded video and audio data
[1065] Output: Data stored in the cloud, suspicious behavior detection results
[1066] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1067] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1068] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1069] [Third embodiment]
[1070] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1071] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1072] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1073] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1074] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1075] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1076] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1077] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1078] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1079] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1080] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1081] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1082] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to respond to visitors efficiently and safely. The following describes specific embodiments of the present invention.
[1083] System Configuration
[1084] The system consists of the following main components:
[1085] 1. Edge AI Camera: Detects visitor movements and captures footage.
[1086] 2. Audio pickup microphone: Captures the voice of visitors.
[1087] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[1088] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1089] 5. Speaker: Communicates the generated voice response to the visitor.
[1090] 6. QR code display device: Displays a QR code for payment and verification.
[1091] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[1092] Explanation of program processing
[1093] The system works by using an edge AI camera and a microphone to capture the visitor's voice and video when they approach and sending it to a cloud server.
[1094] Visitor detection and response
[1095] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1096] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[1097] Speech recognition and intent understanding
[1098] Device: A microphone captures the visitor's voice and sends it to the generative AI model.
[1099] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, "delivery."
[1100] Response generation and display
[1101] Server: Generates an appropriate response based on the information obtained from speech analysis, for example, "You're delivering a package. Please scan the QR code."
[1102] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[1103] QR code payment
[1104] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[1105] Terminal: Confirm that the payment has been completed and repeat the voice message, "Payment has been confirmed. Please place your luggage."
[1106] Recording and Monitoring
[1107] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[1108] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[1109] Specific examples
[1110] Examples:
[1111] This system will be installed on the doors of homes where elderly people live.
[1112] 1. Visitor detection and response
[1113] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[1114] 2. Speech Recognition and Intent Understanding
[1115] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[1116] 3. Response Generation and Display
[1117] The generative AI model generates a response saying, "This is a package delivery. Please scan the QR code," and the response is relayed to the visitor through a speaker. The QR code is then displayed on a digital sign.
[1118] 4. QR code payment
[1119] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[1120] 5. Recording and Monitoring
[1121] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[1122] The above is a specific embodiment for carrying out the present invention, which allows elderly people to safely and comfortably receive visitors and live with peace of mind even when they are away or busy.
[1123] The processing flow will be explained below.
[1124] Step 1:
[1125] Visitor Detection
[1126] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the detection area, the camera captures the visitor's video and transmits the information to the system.
[1127] Step 2:
[1128] Activating the microphone
[1129] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[1130] Step 3:
[1131] Digital human display and greeting
[1132] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1133] Step 4:
[1134] Capture audio data
[1135] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[1136] Step 5:
[1137] Analysis of audio data
[1138] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, such as "delivery."
[1139] Step 6:
[1140] Generating a response
[1141] Server: Generates an appropriate response based on the visitor's intent. For example, "You're about to deliver a package. Please scan the QR code."
[1142] Step 7:
[1143] Sending generated responses and audio output
[1144] Server: Generates and sends the response to the device.
[1145] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[1146] Step 8:
[1147] Displaying the QR code
[1148] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[1149] Step 9:
[1150] Payment confirmation
[1151] User: A visitor scans the QR code with their smartphone and makes a payment.
[1152] Terminal: If the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[1153] Step 10:
[1154] Record of response status
[1155] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[1156] Step 11:
[1157] Data storage and analysis
[1158] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[1159] Step 12:
[1160] User Notification
[1161] Server: Sends visitor information (images, audio, purpose, etc.) to the user and notifies them with a message such as "XX Takkyubin has been delivered."
[1162] Example 1
[1163] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1164] In modern society, responding to visitors needs to be done efficiently and safely. In particular, for elderly people and those living alone, face-to-face interactions with visitors are often a burden. Furthermore, when the home is busy or out of the home, prompt and accurate responses are required, and it is also important to detect suspicious behavior. To solve these issues, a system is needed that can automatically detect visitor movements, understand the visitor's intentions, and provide an appropriate response.
[1165] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1166] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intention, means for outputting the generated response as voice, means for displaying visual information to the visitor, means for recording the visitor's interaction, and means for saving the recorded interaction. This not only enables efficient interaction with visitors, but also enables the visitor's intention to be understood quickly and accurately, enabling appropriate interaction even when the visitor is absent or busy. Furthermore, by detecting suspicious behavior, safety is improved.
[1167] "Means for detecting approaching visitors" refers to devices such as sensors and cameras that detect visitor movement, as well as software that controls them.
[1168] "Means for capturing visitor audio" refers to a microphone for recording the audio made by the visitor, and any equipment or software for appropriately processing that audio data.
[1169] "Means for analyzing voice data to understand visitor intent" refers to algorithms and software for analyzing captured voice data and understanding its content, primarily using generative AI models.
[1170] "Means for generating responses based on understood intent" refers to software and algorithms that understand the visitor's intent and generate appropriate responses based on that content.
[1171] The "means for outputting the generated response as voice" refers to a speaker for outputting the generated response as voice, and a device and software for controlling the speaker.
[1172] "Means for displaying visual information to visitors" refers to devices such as displays and digital signage that provide visual information to visitors, as well as software for controlling them.
[1173] "Means for recording interactions with visitors" refers to devices such as cameras and microphones for recording and recording interactions with visitors and the situation, as well as software for controlling them.
[1174] "Means for storing recorded responses" refers to a storage device for safely storing video and audio data, and software for managing that data.
[1175] System configuration and operation overview
[1176] This system is a digital human-type AI reception system for automatically responding to visitors. The system consists of the following main hardware and software components:
[1177] 1. Edge AI Camera: Detects visitor movements and captures footage.
[1178] 2. Audio pickup microphone: Captures the voice of visitors.
[1179] 3. Digital Signage: Displaying a digital human and responding visually and audibly to visitors.
[1180] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1181] 5. Speaker: Communicates the generated voice response to the visitor.
[1182] 6. QR code display device: Displays a QR code for payment and verification.
[1183] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[1184] System operation details
[1185] Visitor Detection
[1186] Terminal: When the Edge AI camera detects a visitor's movement, the system automatically wakes up. At the same time, the audio pickup microphone activates and captures the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1187] Audio capture and transmission
[1188] Terminal: The visitor's voice (e.g., "This is a parcel delivery") is captured by a microphone and sent to a cloud server. Since the voice data is processed in real time, low-latency data transfer is required.
[1189] Voice analysis and intent understanding
[1190] Server: The cloud server analyzes the received voice data using a generative AI model to understand the visitor's intent. For example, it can extract the intent "delivery" from the voice data "This is a courier delivery."
[1191] Response generation and display
[1192] Server: The generative AI model generates an appropriate response based on the analysis results, for example, "This is a package delivery. Please scan the QR code."
[1193] Terminal: The text response sent from the server is communicated to the visitor using a digital sign and speaker. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the QR code is displayed on the digital sign.
[1194] QR code payment
[1195] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[1196] Terminal: Once the payment is confirmed, the visitor will be notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[1197] Recording and Monitoring
[1198] Terminal: Records and records interactions with visitors and sends the data to a cloud server.
[1199] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[1200] Example operation
[1201] Consider the case where this system is installed on the door of a home where elderly people live.
[1202] 1. Visitor Detection: When a courier approaches the door, the Edge AI camera detects the movement and the system is activated. A digital human appears and asks the courier, "Welcome. How can I help you?"
[1203] 2. Voice capture and transmission: The delivery person responds, "This is a courier delivery," and the voice is captured by the microphone and transmitted to the cloud server.
[1204] 3. Speech analysis and intent understanding: The generative AI model analyzes the voice data and understands the intent of "delivery."
[1205] 4. Response generation and display: The generative AI model generates a response such as "This is a package delivery. Please scan the QR code," plays the audio over the speaker, and displays the QR code on the digital signage.
[1206] 5. QR code payment: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, they will receive a voice message saying, "Payment has been confirmed. Please leave your package."
[1207] 6. Recording and monitoring: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[1208] As a result, this system can respond to visitors efficiently and safely, and can provide safe and secure living support for elderly people and those living alone in their homes.
[1209] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1210] Step 1: Visitor detection
[1211] Terminal: When the Edge AI camera detects a visitor's movement, the system wakes up. The audio pickup microphone activates and prepares to capture the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can we help you?"
[1212] Input: Visitor motion detection (Edge AI camera)
[1213] Output: System startup, digital human display, greeting message output
[1214] Step 2: Capture audio
[1215] Terminal: A microphone captures the visitor's voice, such as "This is a courier delivery," and the voice data is converted into a digital format.
[1216] Input: Visitor's voice
[1217] Output: Digital audio data
[1218] Step 3: Audio transmission and analysis
[1219] Terminal: The captured audio data is sent to the cloud server in real time.
[1220] Server: Analyzes the received voice data using a generative AI model to understand the visitor's intent. Converts the voice data into text and uses natural language processing technology to extract the intent. For example, understand the intent of "delivery" from the voice saying "This is a courier delivery."
[1221] Input: Digital audio data
[1222] Output: Text data containing intent
[1223] Step 4: Response Generation
[1224] Server: The generative AI model generates an appropriate response based on the analysis results, for example, a text response such as "This is a package delivery. Please scan the QR code."
[1225] Input: Text data containing visitor intent
[1226] Output: Response text
[1227] Step 5: Response display and audio output
[1228] Terminal: The response text received from the server is communicated to the visitor via a speaker and digital signage. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the digital signage displays the QR code.
[1229] Input: Response text
[1230] Output: Voice response, QR code display
[1231] Step 6: QR code payment
[1232] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[1233] Terminal: The QR code is scanned to confirm the payment has been completed. Once the payment is complete, the visitor is notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[1234] Input: QR code scan, payment information
[1235] Output: Payment completion notification
[1236] Step 7: Record and monitor
[1237] Terminal: Records and records interactions with visitors and sends the data, including video and audio data, to a cloud server.
[1238] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[1239] Input: Video data, audio data
[1240] Output: Saved data, suspicious behavior notification
[1241] (Application example 1)
[1242] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1243] Modern stores are required to respond to visitors quickly and accurately, but this places a heavy burden on employees, making it difficult for visitors to receive satisfactory service. It is also important to detect suspicious behavior early and respond appropriately. The present invention aims to solve these problems and respond to visitors efficiently and safely. Another objective of the present invention is to enable employees to smoothly confirm visitor responses and improve the efficiency of store operations.
[1244] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1245] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intent, means for generating a response based on the understood intent, means for outputting the generated response as voice, means for recording the interaction with the visitor, means for saving the recorded interaction, and means for an employee wearing smart glasses to visually confirm the response. This allows employees to instantly understand the visitor's intent through the smart glasses and respond promptly and appropriately. Furthermore, recording and saving the interaction makes it easy to review later and detect suspicious behavior.
[1246] A "means for detecting approaching visitors" is a device or method that senses the movement of a visitor and triggers activation of the system.
[1247] A "visitor voice capturing means" is a device or method that collects voice from a visitor and processes it as digital data.
[1248] The "means for analyzing voice data and understanding the visitor's intent" refers to a device or method for analyzing collected voice data and understanding the visitor's requests and questions.
[1249] The "means for generating a response based on the understood intent" is a device or method that automatically generates an appropriate response based on the analysis results.
[1250] The "means for outputting the generated response by voice" is a device or method for transmitting the generated response to the visitor as voice.
[1251] "Means for recording visitor interactions" refers to a device or method for saving the interaction and interaction with visitors as digital data.
[1252] "Means for storing recorded response situations" refers to a device or method for durably storing recorded digital data.
[1253] A "means for visually confirming a response by an employee equipped with smart glasses" is a device or method for visually confirming a response generated through smart glasses worn by an employee.
[1254] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and an embodiment thereof is shown based on an application example using smart glasses worn by employees.
[1255] System configuration:
[1256] The system consists of the following main components:
[1257] 1. Edge AI Camera: When a visitor enters the store, it detects their movement and activates the entire system.
[1258] 2. Audio pickup microphone: Collects visitors' voices and processes them as digital data.
[1259] 3. Smart glasses: Devices worn by employees to visually and audibly confirm the visitor's intentions.
[1260] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1261] 5. Speaker: Outputs the generated response as audio and conveys it to the visitor.
[1262] 6. Cloud server: Stores recorded responses and detects suspicious behavior.
[1263] Explanation of program operation:
[1264] The server controls the edge AI camera, microphone, smart glasses, generative AI model, speaker, and cloud server to enable interaction with visitors. The main processing flow is as follows:
[1265] Visitor Detection:
[1266] When the Edge AI camera detects a visitor's movement, the system is activated, the microphone captures the visitor's voice, and a digital human appears in the smart glasses and asks the visitor, "Welcome. How can I help you?"
[1267] Speech Recognition and Intent Understanding:
[1268] The system captures audio data with a microphone and sends it to a generative AI model, which then analyzes the audio data to understand the visitor's intent.
[1269] For example, if a visitor says, "Please tell me where the new products are," the generative AI model understands the intent "new products" and generates the response, "The new products corner is in the back right."
[1270] Response generation and display:
[1271] The generative AI model generates an appropriate response, which is output through the speaker and also displayed on the smart glasses' display.
[1272] Record of response:
[1273] Visitor interactions are videotaped and stored on a cloud server, and if any suspicious activity is detected, a notification is sent to the administrator.
[1274] Hardware and software used:
[1275] Hardware: Edge AI camera (general security camera), sound pickup microphone (standalone microphone), smart glasses (e.g., smart glasses), speaker.
[1276] Software: speech recognition models (e.g., Google Speech-to-Text API), generative AI models (e.g., OpenAI GPT), and cloud storage (e.g., Google Cloud).
[1277] Example prompt sentence:
[1278] Input speech data: "Please tell me where the new products are located."
[1279] Prompt: "Based on the audio data, analyze the visitor's intent and generate an appropriate response. For example, a response to a question about the location of a new product."
[1280] Thus, the present invention utilizes cutting edge technology such as smart glasses to provide a system that can efficiently and safely serve visitors.
[1281] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1282] Step 1:
[1283] Visitor Detection
[1284] The device's edge AI camera detects the visitor's movements. The motion detection activates the system, which then captures the visitor's video data. The input is the visitor's movements, and the output is the video data and a signal to activate the system.
[1285] Step 2:
[1286] Audio Capture
[1287] The device's microphone captures the visitor's voice, and the collected voice data is sent to the system. The input is the visitor's voice, and the output is digital voice data.
[1288] Step 3:
[1289] Analysis of audio data
[1290] The server's generated AI model analyzes the collected voice data and understands the visitor's intent. The input is digital voice data, and the output is the visitor's intent as an analysis result. A voice recognition model is used for the analysis.
[1291] Step 4:
[1292] Response Generation
[1293] The server's generative AI model generates an appropriate response based on the visitor's intent. The input is the visitor's intent as a result of analysis, and the output is the generated response text. For example, if the question is about the location of new products, the generated response will be "The new products corner is in the back right."
[1294] Step 5:
[1295] Display and speak responses
[1296] The smart glasses on the terminal visually display the generated response, and the speaker outputs the response aloud. The input is the generated response text, and the output is the visual display and audio output. Employees can respond to visitors' questions instantly through the smart glasses.
[1297] Step 6:
[1298] Record of response status
[1299] The device records the conversation with the visitor and sends it to the cloud server. The input is the audio and video data of the conversation with the visitor, and the output is the recorded data sent to the cloud server.
[1300] Step 7:
[1301] Data storage and analysis
[1302] The cloud server stores the recorded response data and analyzes suspicious behavior. The input is video and audio data, and the output is the analysis result, detecting suspicious behavior and notifying the administrator.
[1303] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1304] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to efficiently and safely serve visitors while recognizing their emotions and responding appropriately. A specific embodiment of the present invention will be described below.
[1305] System Configuration
[1306] The system consists of the following main components:
[1307] 1. Edge AI Camera: Detects visitor movements and captures footage.
[1308] 2. Audio pickup microphone: Captures the voice of visitors.
[1309] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[1310] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1311] 5. Emotion Engine: Analyzes audio and video data to recognize visitors' emotions.
[1312] 6. Speaker: Communicates the generated voice response to the visitor.
[1313] 7. QR code display device: Displays a QR code for payment and verification.
[1314] 8. Cloud server: Stores recorded data and analyzes suspicious behavior.
[1315] Explanation of program processing
[1316] The system works by capturing the visitor's voice and video using an edge AI camera and microphone when the visitor approaches, sending the captured video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[1317] Visitor detection and response
[1318] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1319] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[1320] Speech recognition and intent understanding
[1321] Device: A microphone captures the visitor's voice and sends it to a cloud server running a generative AI model.
[1322] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, determining the intent to "deliver."
[1323] emotion recognition
[1324] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone.
[1325] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice, such as "happy," "angry," or "sad."
[1326] Response generation and display
[1327] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[1328] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[1329] QR code payment
[1330] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[1331] Terminal: Confirms that the payment has been completed and repeats the message "Payment has been confirmed. Please place your luggage."
[1332] Recording and Monitoring
[1333] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[1334] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[1335] Specific examples
[1336] Examples:
[1337] This system will be installed on the doors of homes where elderly people live.
[1338] 1. Visitor detection and response
[1339] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[1340] 2. Speech Recognition and Intent Understanding
[1341] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[1342] 3. Emotion recognition
[1343] The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[1344] 4. Response Generation and Display
[1345] The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The response is then played over the speaker to the visitor. The QR code is then displayed on the digital signage.
[1346] 5. QR code payment
[1347] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[1348] 6. Recording and Monitoring
[1349] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[1350] The above is a specific embodiment of the present invention, which allows elderly people to safely and comfortably attend to visitors and furthermore allows responses that take into consideration the feelings of visitors, thereby providing a more user-friendly system.
[1351] The processing flow will be explained below.
[1352] Step 1:
[1353] Visitor Detection
[1354] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the designated detection area, the camera captures the visitor's video and transmits the information to the system.
[1355] Step 2:
[1356] Activating the microphone
[1357] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[1358] Step 3:
[1359] Digital human display and greeting
[1360] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1361] Step 4:
[1362] Capture audio data
[1363] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[1364] Step 5:
[1365] Analysis of audio data
[1366] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, for example, determining the intent "delivery."
[1367] Step 6:
[1368] Emotion recognition
[1369] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone. It recognizes emotions from the visitor's facial expressions and tone of voice.
[1370] Server: The emotion engine uses the analysis results to understand emotions such as "happy," "angry," and "sad."
[1371] Step 7:
[1372] Generating a response
[1373] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[1374] Step 8:
[1375] Sending generated responses and audio output
[1376] Server: Generates and sends the response to the device.
[1377] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[1378] Step 9:
[1379] Displaying the QR code
[1380] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[1381] Step 10:
[1382] Payment confirmation
[1383] User: A visitor scans the QR code with their smartphone and makes a payment.
[1384] Terminal: Once the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[1385] Step 11:
[1386] Record of response status
[1387] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[1388] Step 12:
[1389] Data storage and analysis
[1390] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[1391] Step 13:
[1392] User Notification
[1393] Server: Sends visitor information (image, audio, purpose, emotion, etc.) to the user and notifies them with a message such as "XX delivery has been made."
[1394] Example 2
[1395] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1396] Conventional reception systems have difficulty accurately understanding the visitor's intentions and generating appropriate responses that take their emotions into account. Furthermore, responding without accurately understanding the visitor's intentions and emotions can result in a decline in the quality of service provided to users. Furthermore, the system lacks sufficient functionality to detect suspicious behavior and maintain safety. It is necessary to solve these problems and provide efficient and safe responses to visitors.
[1397] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1398] In this invention, the server includes a means for transmitting voice data to a cloud server, a means for analyzing the voice data in the cloud server to understand the visitor's intention, and a means for performing emotion recognition based on the understood intention. This makes it possible to accurately grasp the visitor's intention and emotion and generate an appropriate response based on that. It is also possible to detect suspicious behavior in real time to ensure safety.
[1399] "Means for detecting approaching visitors" refers to a function that uses a detection device such as a camera or sensor to detect when a visitor approaches a certain location.
[1400] The "means for capturing the visitor's voice" is a function for recording the visitor's speech using a voice input device such as a microphone.
[1401] The "means for transmitting audio data to a cloud server" is a function for transmitting captured audio data to a cloud server via the Internet.
[1402] "Means for analyzing voice data on a cloud server to understand the visitor's intent" refers to a function that uses voice recognition technology on a cloud server to analyze voice data and understand what the visitor is looking for.
[1403] "Means for recognizing emotions based on understood intentions" is a function that understands the visitor's intentions and then determines their emotions from their audio and video data.
[1404] "Means for visually and audibly outputting the generated response" refers to a function that communicates the response generated by the system to the visitor using an output device such as a display or speaker.
[1405] "Means for displaying a QR code as part of a response" refers to the ability to display a QR code on a digital display as a response to a visitor.
[1406] "Means for recording interactions with visitors" refers to the system's ability to record video and audio of interactions with visitors.
[1407] "Means for saving the recorded response status on a cloud server" is a function for saving the recorded video and audio data on a cloud server.
[1408] "Means for selecting the most appropriate response" is a function that determines the response that the system deems most appropriate based on the visitor's intentions and emotions.
[1409] "Means for analyzing and detecting suspicious behavior on a cloud server" refers to a function that analyzes data stored on a cloud server and automatically detects suspicious behavior.
[1410] The present invention is a digital human-type AI reception system that automatically responds when a visitor approaches, accurately understanding the visitor's intentions and emotions and responding optimally based on the results, thereby providing a safe and comfortable experience for the visitor. Specific embodiments of the present invention are described below.
[1411] System Configuration
[1412] The system consists of the following main components:
[1413] 1. Edge AI Camera: A device that detects visitor movements and captures video.
[1414] 2. Audio pickup microphone: A device that captures the voices of visitors.
[1415] 3. Digital Signage: A device that displays a digital human and responds visually and audibly to visitors.
[1416] 4. Generative AI model: Software that runs on a cloud server, analyzes voice data, understands the visitor's intent, and generates appropriate responses.
[1417] 5. Emotion Engine: Software that runs on a cloud server and analyzes audio and video data to recognize visitors' emotions.
[1418] 6. Speaker: A device that transmits the generated voice response to the visitor.
[1419] 7. QR code display device: A device that displays QR codes for payment and verification.
[1420] 8. Cloud Server: A device that stores recorded data and analyzes suspicious behavior.
[1421] Explanation of program processing
[1422] The system works by capturing audio and video of visitors using an edge AI camera and microphone when they approach, sending the captured audio and video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[1423] Visitor detection and response
[1424] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1425] User: The visitor answers with the purpose of the call, such as "This is a courier delivery."
[1426] Speech recognition and intent understanding
[1427] Terminal: A microphone captures the visitor's voice and sends it to a cloud server.
[1428] Server: The generated AI model on the cloud server analyzes the voice data and understands the visitor's intent. For example, it determines whether the intent is "This is a courier delivery."
[1429] emotion recognition
[1430] Device: The video captured by the edge AI camera and the audio captured by the microphone are sent to the emotion engine on the cloud server.
[1431] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. For example, it recognizes the emotion "I'm in a hurry."
[1432] Response generation and display
[1433] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is in a hurry, the server generates a response such as, "You have a package to deliver. If you are in a hurry, please scan the QR code."
[1434] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[1435] QR code payment
[1436] User: Scans the displayed QR code using a smartphone and makes a payment.
[1437] Terminal: Once the payment is confirmed, the terminal will again announce to the visitor, "Payment has been confirmed. Please leave your luggage."
[1438] Recording and Monitoring
[1439] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[1440] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[1441] Specific examples
[1442] Example: This system is installed on the door of a home where elderly people live.
[1443] Visitor detection and response
[1444] Device: A delivery person arrives at the elderly person's home. When the Edge AI camera detects the delivery person, the system is activated. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[1445] User: The delivery person replies, "This is a courier delivery."
[1446] Speech recognition and intent understanding
[1447] Device: A microphone captures the voice and sends it to a cloud server. A generative AI model analyzes the voice and understands the intent of "delivery."
[1448] emotion recognition
[1449] Server: The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[1450] Response generation and display
[1451] Server: The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The speaker plays the message to the visitor. The digital sign displays the QR code.
[1452] QR code payment
[1453] User: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the delivery person is told again by voice, "The payment has been confirmed. Please leave your package."
[1454] Recording and Monitoring
[1455] Device: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[1456] Example prompt sentence:
[1457] An example of a prompt sentence that a visitor can enter into the system using the API is, "This is a courier delivery. It's urgent, so please respond quickly."
[1458] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to safely and comfortably attend to visitors, and furthermore, it is possible to respond in a way that takes into consideration the feelings of visitors, providing a more user-friendly system.
[1459] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1460] Step 1: Visitor detection
[1461] Device: The Edge AI camera detects visitor movement. The camera constantly monitors the surroundings, and when movement is detected, the system automatically activates. Specifically, it identifies human silhouettes and movements and extracts their coordinate data.
[1462] Input: Camera video data
[1463] Output: coordinate data of detected motion and visitor detection trigger signal
[1464] Step 2: Voice capture and digital human response
[1465] Terminal: The microphone begins to capture the visitor's voice. At the same time, a digital human appears on the digital signage and asks, "Welcome. How can I help you?" The captured voice is converted into audio data through a signal processing engine.
[1466] Input: Visitor voice, visitor detection trigger signal
[1467] Output: Converted audio data, activation of digital human representation
[1468] Step 3: Visitor response
[1469] User: The visitor answers with their request, such as "This is a courier delivery." The microphone captures their voice.
[1470] Input: Visitor's response voice
[1471] Output: Captured response audio data
[1472] Step 4: Sending audio data
[1473] Device: The device sends the captured audio data to a cloud server, where it is converted into packets and transmitted using a secure protocol.
[1474] Input: Captured response audio data
[1475] Output: Send audio data to cloud server
[1476] Step 5: Voice data analysis and intent understanding
[1477] Server: The AI model generated on the cloud server analyzes the voice data and understands the visitor's intent. The voice data is converted into text data through a natural language processing engine, and the intent is extracted from the text data. For example, the intent "delivery" is understood from the phrase "This is a courier delivery."
[1478] Input: Audio data sent to the cloud server
[1479] Output: Text data about the visitor's intent (e.g., "delivery")
[1480] Step 6: Sending Emotion Data
[1481] Device: Video captured by the edge AI camera and audio captured by the microphone are sent to the emotion engine on the cloud server. The video and audio data are integrated and sent as a single data packet.
[1482] Input: Video and audio data
[1483] Output: Sending sentiment analysis data to the cloud server
[1484] Step 7: Sentiment Data Analysis
[1485] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. It extracts facial features from video data and analyzes tone characteristics from audio data. For example, it recognizes emotions such as "I'm in a hurry."
[1486] Input: Video and audio data sent to the cloud server
[1487] Output: Data about visitor sentiment (e.g., "I'm in a hurry")
[1488] Step 8: Response Generation
[1489] Server: Generates appropriate responses based on information obtained through speech analysis and emotion recognition. The generative AI model generates the optimal answer based on the visitor's intent and emotion data. For example, it generates a response such as, "You're delivering a package. If it's urgent, please scan the QR code."
[1490] Input: Visitor intent data, sentiment data
[1491] Output: Text data of the response
[1492] Step 9: Response display and communication
[1493] Terminal: The generated response is transmitted to the visitor via a speaker. A QR code is displayed on the digital signage. The response text is converted into speech using a speech synthesis engine and output from the speaker.
[1494] Input: Text data of response content
[1495] Output: Audio and visual response (QR code display)
[1496] Step 10: Scan the QR code and pay
[1497] User: Scans the displayed QR code using a smartphone and makes the payment. Once the payment is completed through the payment app, a notification is automatically sent to the system.
[1498] Input: QR code
[1499] Output: Payment completion notification
[1500] Step 11: Payment confirmation and instructions
[1501] Terminal: Confirm that the payment has been completed and tell the visitor again by voice, "Payment has been confirmed. Please leave your luggage."
[1502] Input: Payment completion notification
[1503] Output: Voice notification of payment completion
[1504] Step 12: Data recording
[1505] Device: Video and audio recordings are made during the call and sent to a cloud server. The recorded data is encrypted and stored securely.
[1506] Input: Video and audio data during the call
[1507] Output: Send recorded data to cloud server
[1508] Step 13: Detect and notify suspicious behavior
[1509] Server: Analyzes stored data and detects suspicious behavior. It uses algorithms to identify anomalous patterns of behavior and notifies the user if any are detected.
[1510] Input: Stored video and audio data
[1511] Output: Suspicious behavior detected notification
[1512] (Application example 2)
[1513] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1514] Currently, there are many systems that automatically respond to visitors, but few can recognize the visitor's emotions and respond appropriately. Furthermore, only a limited number of systems have the functionality to record the results of the response and detect suspicious behavior. Therefore, there is a demand for a system that can respond to visitors efficiently and safely, and that generates responses that take the visitor's emotions into consideration.
[1515] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for detecting when a visitor is approaching, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotions, means for outputting the generated response by voice and visually displaying it, means for recording the interaction with the visitor, and means for saving the recorded interaction and analyzing it as needed. This enables appropriate interaction that takes the visitor's emotions into consideration, thereby achieving safe and efficient visitor interaction.
[1516] "Means for detecting approaching visitors" refers to devices or software that detect the movement of visitors and activate the system.
[1517] "Means for capturing visitor audio" means a microphone or recording device used to collect and record audio emitted by a visitor.
[1518] "Means for analyzing captured voice data to understand the visitor's intent" refers to algorithms or software that analyzes voice data using voice recognition technology and understands the visitor's intent and requests.
[1519] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on collected data.
[1520] "Means for outputting the generated response audibly and displaying it visually" refers to equipment or software for outputting the generated response audibly through a speaker and displaying it on a display such as digital signage.
[1521] "Visitor interaction recording means" means any device or software that records visitor interactions in video or audio format.
[1522] "Means for storing the recorded response status and analyzing it as needed" refers to software and hardware for storing and analyzing the recorded data in a database or cloud server.
[1523] "Means for displaying QR codes" refers to devices or software that generate QR codes that encode specific information and display them on a display.
[1524] "Means for detecting and responding to suspicious behavior" refers to algorithms or software that analyze recorded data, detect abnormalities or suspicious behavior, and take appropriate action.
[1525] The present invention is a system for automatically responding to visitors, recognizing their emotions, and responding appropriately. This system includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotion, means for outputting the generated response as voice and visually displaying it, means for recording the visitor interaction, and means for saving the recorded interaction and analyzing it as needed.
[1526] System Programming and Processing
[1527] The operation of this system is described as follows: This system uses the following hardware and software:
[1528] Hardware: Smartphone (camera, microphone), server, display, speaker
[1529] Software: OpenCV (image processing), SpeechRecognition (voice recognition), Transformers (generative AI model), TensorFlow (emotion recognition), qrcode (QR code generation), Firebase (data storage and analysis)
[1530] How to detect approaching visitors:
[1531] The system uses the smartphone camera to detect the movement of visitors. For example, when the smartphone camera recognizes a visitor, the system is activated.
[1532] To capture visitor audio:
[1533] The microphone on the smartphone is used to capture the visitor's voice. For example, if a visitor says, "I want to see new products," the microphone will collect that voice.
[1534] How to analyze captured audio data to understand visitor intent:
[1535] The collected voice data is converted into text data using the SpeechRecognition library. A generative AI model is then used to understand the visitor's intent. For example, a speech that says "I want to see new products" is understood to mean "I'm looking for product information."
[1536] Measures including generative AI models that generate responses based on the understood intent and sentiment of the visitor:
[1537] A response is generated based on the understood intent and the visitor's emotion (e.g., "excited") analyzed using an emotion recognition model using TensorFlow. Transformers is used as the generative AI model to input appropriate prompt sentences. For example, if the visitor is excited, the response generated will be, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[1538] A way to both speak and visually display the generated response:
[1539] The generated response is output as audio through the smartphone's speaker and displayed visually on the display, along with a QR code. For example, the QR code may be displayed along with a voice message saying, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[1540] How we record visitor interactions:
[1541] Visitor interactions are recorded in video and audio format, for example when a visitor scans a QR code.
[1542] A means to store recorded interactions and analyze them as needed:
[1543] The recorded data is stored in the cloud using Firebase and analyzed as needed to detect suspicious behavior. For example, if abnormal movements or behavior are detected, an administrator will be notified.
[1544] Examples and prompts
[1545] As a concrete example, consider a system placed at the entrance of a physical store. When a visitor approaches the store entrance, the system activates and analyzes the visitor's facial expressions and voice to respond. For example, along with a voice response such as "Welcome. How can I help you?", a response may be generated that instructs the visitor, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[1546] Prompt Sentence Examples
[1547] "The visitor's sentiment is 'Excited' and they say: 'I want to see new products.' Generate an appropriate response. Visitor: 'I want to see new products.' Example of an appropriate response: 'Here is a list of new products. Choose your favorite and scan the QR code.'"
[1548] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1549] Step 1:
[1550] A means (device) of detecting approaching visitors
[1551] The system uses the smartphone camera to detect the movement of visitors. Specifically, it uses the OpenCV library to analyze the camera footage and perform motion detection. When motion is detected, the system is activated.
[1552] Input: Camera video data
[1553] Output: A signal that a visitor is approaching
[1554] Step 2:
[1555] A means (device) for capturing visitor audio
[1556] It uses the device's microphone to capture the visitor's voice, specifically by using the SpeechRecognition library to collect the voice data.
[1557] Input: Visitor utterance
[1558] Output: Audio data
[1559] Step 3:
[1560] A means of analyzing voice data and understanding the visitor's intent (server)
[1561] The captured voice data is sent to the server, where it is converted into text using the SpeechRecognition library, and then generative AI models (Transformers) are used to understand the visitor's intent.
[1562] Input: Audio data
[1563] Output: Text data containing visitor intent
[1564] Step 4:
[1565] A means of analyzing visitor sentiment (server)
[1566] Using video data captured on the server, we analyze visitors' emotions using a TensorFlow model, specifically identifying emotions from facial expressions and tone of voice.
[1567] Input: Video and audio data
[1568] Output: Visitor sentiment data
[1569] Step 5:
[1570] A means (server) to generate responses based on the understood intent and emotions
[1571] Using generative AI models (Transformers), it generates appropriate responses based on the visitor's intent and emotions. For example, if the visitor is excited, it might generate a response like, "Here's a list of new products. Choose your favorite and scan the QR code."
[1572] Input: Text data containing intent, emotion data
[1573] Output: The generated response text
[1574] Step 6:
[1575] A means (terminal) to output the generated response as voice and display it visually
[1576] The generated response text is synthesized into speech and output as voice through the smartphone speaker. The response text including a QR code is also displayed on the display. The QR code is generated using the qrcode library.
[1577] Input: Generated response text
[1578] Output: Audio output, visual display with QR code
[1579] Step 7:
[1580] A means (terminal) for recording visitor interactions
[1581] Record visitor interactions in video and audio format using the smartphone camera and microphone.
[1582] Input: Video and audio when greeting a visitor
[1583] Output: Recorded video and audio data
[1584] Step 8:
[1585] A means (server) to store the recorded response status and analyze it as needed
[1586] The recorded video and audio data is stored on the Firebase cloud server. If necessary, the stored data is analyzed, and if any suspicious activity is detected, an administrator is notified.
[1587] Input: Recorded video and audio data
[1588] Output: Data stored in the cloud, suspicious behavior detection results
[1589] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1590] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1591] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1592] [Fourth embodiment]
[1593] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1594] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1595] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1596] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1597] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1598] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1599] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1600] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1601] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1602] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1603] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1604] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1605] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1606] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to respond to visitors efficiently and safely. The following describes specific embodiments of the present invention.
[1607] System Configuration
[1608] The system consists of the following main components:
[1609] 1. Edge AI Camera: Detects visitor movements and captures footage.
[1610] 2. Audio pickup microphone: Captures the voice of visitors.
[1611] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[1612] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1613] 5. Speaker: Communicates the generated voice response to the visitor.
[1614] 6. QR code display device: Displays a QR code for payment and verification.
[1615] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[1616] Explanation of program processing
[1617] The system works by using an edge AI camera and a microphone to capture the visitor's voice and video when they approach and sending it to a cloud server.
[1618] Visitor detection and response
[1619] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1620] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[1621] Speech recognition and intent understanding
[1622] Device: A microphone captures the visitor's voice and sends it to the generative AI model.
[1623] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, "delivery."
[1624] Response generation and display
[1625] Server: Generates an appropriate response based on the information obtained from speech analysis, for example, "You're delivering a package. Please scan the QR code."
[1626] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[1627] QR code payment
[1628] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[1629] Terminal: Confirm that the payment has been completed and repeat the voice message, "Payment has been confirmed. Please place your luggage."
[1630] Recording and Monitoring
[1631] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[1632] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[1633] Specific examples
[1634] Examples:
[1635] This system will be installed on the doors of homes where elderly people live.
[1636] 1. Visitor detection and response
[1637] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[1638] 2. Speech Recognition and Intent Understanding
[1639] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[1640] 3. Response Generation and Display
[1641] The generative AI model generates a response saying, "This is a package delivery. Please scan the QR code," and the response is relayed to the visitor through a speaker. The QR code is then displayed on a digital sign.
[1642] 4. QR code payment
[1643] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[1644] 5. Recording and Monitoring
[1645] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[1646] The above is a specific embodiment for carrying out the present invention, which allows elderly people to safely and comfortably receive visitors and live with peace of mind even when they are away or busy.
[1647] The processing flow will be explained below.
[1648] Step 1:
[1649] Visitor Detection
[1650] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the detection area, the camera captures the visitor's video and transmits the information to the system.
[1651] Step 2:
[1652] Activating the microphone
[1653] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[1654] Step 3:
[1655] Digital human display and greeting
[1656] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1657] Step 4:
[1658] Capture audio data
[1659] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[1660] Step 5:
[1661] Analysis of audio data
[1662] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, such as "delivery."
[1663] Step 6:
[1664] Generating a response
[1665] Server: Generates an appropriate response based on the visitor's intent. For example, "You're about to deliver a package. Please scan the QR code."
[1666] Step 7:
[1667] Sending generated responses and audio output
[1668] Server: Generates and sends the response to the device.
[1669] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[1670] Step 8:
[1671] Displaying the QR code
[1672] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[1673] Step 9:
[1674] Payment confirmation
[1675] User: A visitor scans the QR code with their smartphone and makes a payment.
[1676] Terminal: If the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[1677] Step 10:
[1678] Record of response status
[1679] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[1680] Step 11:
[1681] Data storage and analysis
[1682] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[1683] Step 12:
[1684] User Notification
[1685] Server: Sends visitor information (images, audio, purpose, etc.) to the user and notifies them with a message such as "XX Takkyubin has been delivered."
[1686] Example 1
[1687] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1688] In modern society, responding to visitors needs to be done efficiently and safely. In particular, for elderly people and those living alone, face-to-face interactions with visitors are often a burden. Furthermore, when the home is busy or out of the home, prompt and accurate responses are required, and it is also important to detect suspicious behavior. To solve these issues, a system is needed that can automatically detect visitor movements, understand the visitor's intentions, and provide an appropriate response.
[1689] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1690] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intention, means for outputting the generated response as voice, means for displaying visual information to the visitor, means for recording the visitor's interaction, and means for saving the recorded interaction. This not only enables efficient interaction with visitors, but also enables the visitor's intention to be understood quickly and accurately, enabling appropriate interaction even when the visitor is absent or busy. Furthermore, by detecting suspicious behavior, safety is improved.
[1691] "Means for detecting approaching visitors" refers to devices such as sensors and cameras that detect visitor movement, as well as software that controls them.
[1692] "Means for capturing visitor audio" refers to a microphone for recording the audio made by the visitor, and any equipment or software for appropriately processing that audio data.
[1693] "Means for analyzing voice data to understand visitor intent" refers to algorithms and software for analyzing captured voice data and understanding its content, primarily using generative AI models.
[1694] "Means for generating responses based on understood intent" refers to software and algorithms that understand the visitor's intent and generate appropriate responses based on that content.
[1695] The "means for outputting the generated response as voice" refers to a speaker for outputting the generated response as voice, and a device and software for controlling the speaker.
[1696] "Means for displaying visual information to visitors" refers to devices such as displays and digital signage that provide visual information to visitors, as well as software for controlling them.
[1697] "Means for recording interactions with visitors" refers to devices such as cameras and microphones for recording and recording interactions with visitors and the situation, as well as software for controlling them.
[1698] "Means for storing recorded responses" refers to a storage device for safely storing video and audio data, and software for managing that data.
[1699] System configuration and operation overview
[1700] This system is a digital human-type AI reception system for automatically responding to visitors. The system consists of the following main hardware and software components:
[1701] 1. Edge AI Camera: Detects visitor movements and captures footage.
[1702] 2. Audio pickup microphone: Captures the voice of visitors.
[1703] 3. Digital Signage: Displaying a digital human and responding visually and audibly to visitors.
[1704] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1705] 5. Speaker: Communicates the generated voice response to the visitor.
[1706] 6. QR code display device: Displays a QR code for payment and verification.
[1707] 7. Cloud server: Stores recorded data and analyzes suspicious behavior.
[1708] System operation details
[1709] Visitor Detection
[1710] Terminal: When the Edge AI camera detects a visitor's movement, the system automatically wakes up. At the same time, the audio pickup microphone activates and captures the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1711] Audio capture and transmission
[1712] Terminal: The visitor's voice (e.g., "This is a parcel delivery") is captured by a microphone and sent to a cloud server. Since the voice data is processed in real time, low-latency data transfer is required.
[1713] Voice analysis and intent understanding
[1714] Server: The cloud server analyzes the received voice data using a generative AI model to understand the visitor's intent. For example, it can extract the intent "delivery" from the voice data "This is a courier delivery."
[1715] Response generation and display
[1716] Server: The generative AI model generates an appropriate response based on the analysis results, for example, "This is a package delivery. Please scan the QR code."
[1717] Terminal: The text response sent from the server is communicated to the visitor using a digital sign and speaker. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the QR code is displayed on the digital sign.
[1718] QR code payment
[1719] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[1720] Terminal: Once the payment is confirmed, the visitor will be notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[1721] Recording and Monitoring
[1722] Terminal: Records and records interactions with visitors and sends the data to a cloud server.
[1723] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[1724] Example operation
[1725] Consider the case where this system is installed on the door of a home where elderly people live.
[1726] 1. Visitor Detection: When a courier approaches the door, the Edge AI camera detects the movement and the system is activated. A digital human appears and asks the courier, "Welcome. How can I help you?"
[1727] 2. Voice capture and transmission: The delivery person responds, "This is a courier delivery," and the voice is captured by the microphone and transmitted to the cloud server.
[1728] 3. Speech analysis and intent understanding: The generative AI model analyzes the voice data and understands the intent of "delivery."
[1729] 4. Response generation and display: The generative AI model generates a response such as "This is a package delivery. Please scan the QR code," plays the audio over the speaker, and displays the QR code on the digital signage.
[1730] 5. QR code payment: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, they will receive a voice message saying, "Payment has been confirmed. Please leave your package."
[1731] 6. Recording and monitoring: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[1732] As a result, this system can respond to visitors efficiently and safely, and can provide safe and secure living support for elderly people and those living alone in their homes.
[1733] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1734] Step 1: Visitor detection
[1735] Terminal: When the Edge AI camera detects a visitor's movement, the system wakes up. The audio pickup microphone activates and prepares to capture the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can we help you?"
[1736] Input: Visitor motion detection (Edge AI camera)
[1737] Output: System startup, digital human display, greeting message output
[1738] Step 2: Capture audio
[1739] Terminal: A microphone captures the visitor's voice, such as "This is a courier delivery," and the voice data is converted into a digital format.
[1740] Input: Visitor's voice
[1741] Output: Digital audio data
[1742] Step 3: Audio transmission and analysis
[1743] Terminal: The captured audio data is sent to the cloud server in real time.
[1744] Server: Analyzes the received voice data using a generative AI model to understand the visitor's intent. Converts the voice data into text and uses natural language processing technology to extract the intent. For example, understand the intent of "delivery" from the voice saying "This is a courier delivery."
[1745] Input: Digital audio data
[1746] Output: Text data containing intent
[1747] Step 4: Response Generation
[1748] Server: The generative AI model generates an appropriate response based on the analysis results, for example, a text response such as "This is a package delivery. Please scan the QR code."
[1749] Input: Text data containing visitor intent
[1750] Output: Response text
[1751] Step 5: Response display and audio output
[1752] Terminal: The response text received from the server is communicated to the visitor via a speaker and digital signage. The speaker plays a voice message saying, "This is a parcel delivery. Please scan the QR code," and the digital signage displays the QR code.
[1753] Input: Response text
[1754] Output: Voice response, QR code display
[1755] Step 6: QR code payment
[1756] Visitors: Visitors use their smartphones to scan the QR code and make payments.
[1757] Terminal: The QR code is scanned to confirm the payment has been completed. Once the payment is complete, the visitor is notified again with a voice message saying, "Payment has been confirmed. Please leave your luggage."
[1758] Input: QR code scan, payment information
[1759] Output: Payment completion notification
[1760] Step 7: Record and monitor
[1761] Terminal: Records and records interactions with visitors and sends the data, including video and audio data, to a cloud server.
[1762] Server: The cloud server periodically analyzes the stored data and notifies the user if any suspicious activity is detected. Notifications are sent via the user's smartphone app.
[1763] Input: Video data, audio data
[1764] Output: Saved data, suspicious behavior notification
[1765] (Application example 1)
[1766] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1767] Modern stores are required to respond to visitors quickly and accurately, but this places a heavy burden on employees, making it difficult for visitors to receive satisfactory service. It is also important to detect suspicious behavior early and respond appropriately. The present invention aims to solve these problems and respond to visitors efficiently and safely. Another objective of the present invention is to enable employees to smoothly confirm visitor responses and improve the efficiency of store operations.
[1768] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1769] In this invention, the server includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the voice data to understand the visitor's intent, means for generating a response based on the understood intent, means for outputting the generated response as voice, means for recording the interaction with the visitor, means for saving the recorded interaction, and means for an employee wearing smart glasses to visually confirm the response. This allows employees to instantly understand the visitor's intent through the smart glasses and respond promptly and appropriately. Furthermore, recording and saving the interaction makes it easy to review later and detect suspicious behavior.
[1770] A "means for detecting approaching visitors" is a device or method that senses the movement of a visitor and triggers activation of the system.
[1771] A "visitor voice capturing means" is a device or method that collects voice from a visitor and processes it as digital data.
[1772] The "means for analyzing voice data and understanding the visitor's intent" refers to a device or method for analyzing collected voice data and understanding the visitor's requests and questions.
[1773] The "means for generating a response based on the understood intent" is a device or method that automatically generates an appropriate response based on the analysis results.
[1774] The "means for outputting the generated response by voice" is a device or method for transmitting the generated response to the visitor as voice.
[1775] "Means for recording visitor interactions" refers to a device or method for saving the interaction and interaction with visitors as digital data.
[1776] "Means for storing recorded response situations" refers to a device or method for durably storing recorded digital data.
[1777] A "means for visually confirming a response by an employee equipped with smart glasses" is a device or method for visually confirming a response generated through smart glasses worn by an employee.
[1778] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and an embodiment thereof is shown based on an application example using smart glasses worn by employees.
[1779] System configuration:
[1780] The system consists of the following main components:
[1781] 1. Edge AI Camera: When a visitor enters the store, it detects their movement and activates the entire system.
[1782] 2. Audio pickup microphone: Collects visitors' voices and processes them as digital data.
[1783] 3. Smart glasses: Devices worn by employees to visually and audibly confirm the visitor's intentions.
[1784] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1785] 5. Speaker: Outputs the generated response as audio and conveys it to the visitor.
[1786] 6. Cloud server: Stores recorded responses and detects suspicious behavior.
[1787] Explanation of program operation:
[1788] The server controls the edge AI camera, microphone, smart glasses, generative AI model, speaker, and cloud server to enable interaction with visitors. The main processing flow is as follows:
[1789] Visitor Detection:
[1790] When the Edge AI camera detects a visitor's movement, the system is activated, the microphone captures the visitor's voice, and a digital human appears in the smart glasses and asks the visitor, "Welcome. How can I help you?"
[1791] Speech Recognition and Intent Understanding:
[1792] The system captures audio data with a microphone and sends it to a generative AI model, which then analyzes the audio data to understand the visitor's intent.
[1793] For example, if a visitor says, "Please tell me where the new products are," the generative AI model understands the intent "new products" and generates the response, "The new products corner is in the back right."
[1794] Response generation and display:
[1795] The generative AI model generates an appropriate response, which is output through the speaker and also displayed on the smart glasses' display.
[1796] Record of response:
[1797] Visitor interactions are videotaped and stored on a cloud server, and if any suspicious activity is detected, a notification is sent to the administrator.
[1798] Hardware and software used:
[1799] Hardware: Edge AI camera (general security camera), sound pickup microphone (standalone microphone), smart glasses (e.g., smart glasses), speaker.
[1800] Software: speech recognition models (e.g., Google Speech-to-Text API), generative AI models (e.g., OpenAI GPT), and cloud storage (e.g., Google Cloud).
[1801] Example prompt sentence:
[1802] Input speech data: "Please tell me where the new products are located."
[1803] Prompt: "Based on the audio data, analyze the visitor's intent and generate an appropriate response. For example, a response to a question about the location of a new product."
[1804] Thus, the present invention utilizes cutting edge technology such as smart glasses to provide a system that can efficiently and safely serve visitors.
[1805] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1806] Step 1:
[1807] Visitor Detection
[1808] The device's edge AI camera detects the visitor's movements. The motion detection activates the system, which then captures the visitor's video data. The input is the visitor's movements, and the output is the video data and a signal to activate the system.
[1809] Step 2:
[1810] Audio Capture
[1811] The device's microphone captures the visitor's voice, and the collected voice data is sent to the system. The input is the visitor's voice, and the output is digital voice data.
[1812] Step 3:
[1813] Analysis of audio data
[1814] The server's generated AI model analyzes the collected voice data and understands the visitor's intent. The input is digital voice data, and the output is the visitor's intent as an analysis result. A voice recognition model is used for the analysis.
[1815] Step 4:
[1816] Response Generation
[1817] The server's generative AI model generates an appropriate response based on the visitor's intent. The input is the visitor's intent as a result of analysis, and the output is the generated response text. For example, if the question is about the location of new products, the generated response will be "The new products corner is in the back right."
[1818] Step 5:
[1819] Display and speak responses
[1820] The smart glasses on the terminal visually display the generated response, and the speaker outputs the response aloud. The input is the generated response text, and the output is the visual display and audio output. Employees can respond to visitors' questions instantly through the smart glasses.
[1821] Step 6:
[1822] Record of response status
[1823] The device records the conversation with the visitor and sends it to the cloud server. The input is the audio and video data of the conversation with the visitor, and the output is the recorded data sent to the cloud server.
[1824] Step 7:
[1825] Data storage and analysis
[1826] The cloud server stores the recorded response data and analyzes suspicious behavior. The input is video and audio data, and the output is the analysis result, detecting suspicious behavior and notifying the administrator.
[1827] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1828] The present invention relates to a digital human-type AI reception system that automatically responds to visitors when they approach, and aims to efficiently and safely serve visitors while recognizing their emotions and responding appropriately. A specific embodiment of the present invention will be described below.
[1829] System Configuration
[1830] The system consists of the following main components:
[1831] 1. Edge AI Camera: Detects visitor movements and captures footage.
[1832] 2. Audio pickup microphone: Captures the voice of visitors.
[1833] 3. Digital Signage: Display a digital human and respond visually and audibly to visitors.
[1834] 4. Generative AI model: Analyzes voice data, understands visitor intent, and generates appropriate responses.
[1835] 5. Emotion Engine: Analyzes audio and video data to recognize visitors' emotions.
[1836] 6. Speaker: Communicates the generated voice response to the visitor.
[1837] 7. QR code display device: Displays a QR code for payment and verification.
[1838] 8. Cloud server: Stores recorded data and analyzes suspicious behavior.
[1839] Explanation of program processing
[1840] The system works by capturing the visitor's voice and video using an edge AI camera and microphone when the visitor approaches, sending the captured video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[1841] Visitor detection and response
[1842] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1843] Visitor: The visitor will answer the call, such as "This is a courier delivery."
[1844] Speech recognition and intent understanding
[1845] Device: A microphone captures the visitor's voice and sends it to a cloud server running a generative AI model.
[1846] Server: A generative AI model analyzes the voice data and understands the visitor's intent, for example, determining the intent to "deliver."
[1847] emotion recognition
[1848] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone.
[1849] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice, such as "happy," "angry," or "sad."
[1850] Response generation and display
[1851] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[1852] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[1853] QR code payment
[1854] Visitors: Scan the displayed QR code using their smartphone to make a payment.
[1855] Terminal: Confirms that the payment has been completed and repeats the message "Payment has been confirmed. Please place your luggage."
[1856] Recording and Monitoring
[1857] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[1858] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[1859] Specific examples
[1860] Examples:
[1861] This system will be installed on the doors of homes where elderly people live.
[1862] 1. Visitor detection and response
[1863] A delivery person arrives at an elderly person's home. The Edge AI camera detects the delivery person and activates the system. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[1864] 2. Speech Recognition and Intent Understanding
[1865] The delivery person responds, "This is a courier delivery." The microphone captures the voice and sends it to a cloud server. The generative AI model analyzes the voice and understands the intent of "delivery."
[1866] 3. Emotion recognition
[1867] The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[1868] 4. Response Generation and Display
[1869] The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The response is then played over the speaker to the visitor. The QR code is then displayed on the digital signage.
[1870] 5. QR code payment
[1871] The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the customer is told again by voice, "The payment has been confirmed. Please leave your package."
[1872] 6. Recording and Monitoring
[1873] Video and audio recordings of the call are saved on a cloud server, and if any suspicious activity is detected, a notification is sent to the user's smartphone.
[1874] The above is a specific embodiment of the present invention, which allows elderly people to safely and comfortably attend to visitors and furthermore allows responses that take into consideration the feelings of visitors, thereby providing a more user-friendly system.
[1875] The processing flow will be explained below.
[1876] Step 1:
[1877] Visitor Detection
[1878] Terminal: When the Edge AI camera detects motion, the entire system is automatically activated. When a visitor enters the designated detection area, the camera captures the visitor's video and transmits the information to the system.
[1879] Step 2:
[1880] Activating the microphone
[1881] Terminal: The microphone activates and prepares to capture the visitor's voice. If the visitor speaks, the voice is captured as data.
[1882] Step 3:
[1883] Digital human display and greeting
[1884] Terminal: A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1885] Step 4:
[1886] Capture audio data
[1887] Device: A microphone captures the visitor's voice and sends the data to a cloud server where the generative AI model runs.
[1888] Step 5:
[1889] Analysis of audio data
[1890] Server: The generative AI model analyzes the transmitted voice data and understands the visitor's intent, for example, determining the intent "delivery."
[1891] Step 6:
[1892] Emotion recognition
[1893] Device: The emotion engine analyzes the video captured by the edge AI camera and the audio captured by the microphone. It recognizes emotions from the visitor's facial expressions and tone of voice.
[1894] Server: The emotion engine uses the analysis results to understand emotions such as "happy," "angry," and "sad."
[1895] Step 7:
[1896] Generating a response
[1897] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is angry, the server generates a response such as "We apologize for the inconvenience. Please scan the QR code."
[1898] Step 8:
[1899] Sending generated responses and audio output
[1900] Server: Generates and sends the response to the device.
[1901] Terminal: The generated response is audibly transmitted to the visitor through a speaker.
[1902] Step 9:
[1903] Displaying the QR code
[1904] Terminal: Display a QR code on the digital signage. Visitors can scan the QR code with their smartphones to make payments.
[1905] Step 10:
[1906] Payment confirmation
[1907] User: A visitor scans the QR code with their smartphone and makes a payment.
[1908] Terminal: Once the payment is successful, the visitor will hear a voice message saying, "Payment confirmed. Please place your luggage."
[1909] Step 11:
[1910] Record of response status
[1911] Device: Uses a camera and microphone to record interactions with visitors (video and audio) and send them to a cloud server.
[1912] Step 12:
[1913] Data storage and analysis
[1914] Server: Stores audio and video data in the cloud, analyzes suspicious behavior as needed, and notifies the user if any suspicious behavior is detected.
[1915] Step 13:
[1916] User Notification
[1917] Server: Sends visitor information (image, audio, purpose, emotion, etc.) to the user and notifies them with a message such as "XX delivery has been made."
[1918] Example 2
[1919] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1920] Conventional reception systems have difficulty accurately understanding the visitor's intentions and generating appropriate responses that take their emotions into account. Furthermore, responding without accurately understanding the visitor's intentions and emotions can result in a decline in the quality of service provided to users. Furthermore, the system lacks sufficient functionality to detect suspicious behavior and maintain safety. It is necessary to solve these problems and provide efficient and safe responses to visitors.
[1921] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1922] In this invention, the server includes a means for transmitting voice data to a cloud server, a means for analyzing the voice data in the cloud server to understand the visitor's intention, and a means for performing emotion recognition based on the understood intention. This makes it possible to accurately grasp the visitor's intention and emotion and generate an appropriate response based on that. It is also possible to detect suspicious behavior in real time to ensure safety.
[1923] "Means for detecting approaching visitors" refers to a function that uses a detection device such as a camera or sensor to detect when a visitor approaches a certain location.
[1924] The "means for capturing the visitor's voice" is a function for recording the visitor's speech using a voice input device such as a microphone.
[1925] The "means for transmitting audio data to a cloud server" is a function for transmitting captured audio data to a cloud server via the Internet.
[1926] "Means for analyzing voice data on a cloud server to understand the visitor's intent" refers to a function that uses voice recognition technology on a cloud server to analyze voice data and understand what the visitor is looking for.
[1927] "Means for recognizing emotions based on understood intentions" is a function that understands the visitor's intentions and then determines their emotions from their audio and video data.
[1928] "Means for visually and audibly outputting the generated response" refers to a function that communicates the response generated by the system to the visitor using an output device such as a display or speaker.
[1929] "Means for displaying a QR code as part of a response" refers to the ability to display a QR code on a digital display as a response to a visitor.
[1930] "Means for recording interactions with visitors" refers to the system's ability to record video and audio of interactions with visitors.
[1931] "Means for saving the recorded response status on a cloud server" is a function for saving the recorded video and audio data on a cloud server.
[1932] "Means for selecting the most appropriate response" is a function that determines the response that the system deems most appropriate based on the visitor's intentions and emotions.
[1933] "Means for analyzing and detecting suspicious behavior on a cloud server" refers to a function that analyzes data stored on a cloud server and automatically detects suspicious behavior.
[1934] The present invention is a digital human-type AI reception system that automatically responds when a visitor approaches, accurately understanding the visitor's intentions and emotions and responding optimally based on the results, thereby providing a safe and comfortable experience for the visitor. Specific embodiments of the present invention are described below.
[1935] System Configuration
[1936] The system consists of the following main components:
[1937] 1. Edge AI Camera: A device that detects visitor movements and captures video.
[1938] 2. Audio pickup microphone: A device that captures the voices of visitors.
[1939] 3. Digital Signage: A device that displays a digital human and responds visually and audibly to visitors.
[1940] 4. Generative AI model: Software that runs on a cloud server, analyzes voice data, understands the visitor's intent, and generates appropriate responses.
[1941] 5. Emotion Engine: Software that runs on a cloud server and analyzes audio and video data to recognize visitors' emotions.
[1942] 6. Speaker: A device that transmits the generated voice response to the visitor.
[1943] 7. QR code display device: A device that displays QR codes for payment and verification.
[1944] 8. Cloud Server: A device that stores recorded data and analyzes suspicious behavior.
[1945] Explanation of program processing
[1946] The system works by capturing audio and video of visitors using an edge AI camera and microphone when they approach, sending the captured audio and video to a cloud server, and then using an emotion engine to recognize the visitor's emotions and respond appropriately.
[1947] Visitor detection and response
[1948] Terminal: When the Edge AI camera detects motion, the system is activated and the microphone begins capturing the visitor's voice. A digital human appears on the digital signage and asks the visitor, "Welcome. How can I help you?"
[1949] User: The visitor answers with the purpose of the call, such as "This is a courier delivery."
[1950] Speech recognition and intent understanding
[1951] Terminal: A microphone captures the visitor's voice and sends it to a cloud server.
[1952] Server: The generated AI model on the cloud server analyzes the voice data and understands the visitor's intent. For example, it determines whether the intent is "This is a courier delivery."
[1953] emotion recognition
[1954] Device: The video captured by the edge AI camera and the audio captured by the microphone are sent to the emotion engine on the cloud server.
[1955] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. For example, it recognizes the emotion "I'm in a hurry."
[1956] Response generation and display
[1957] Server: Generates an appropriate response based on information obtained through speech analysis and emotion recognition. For example, if the visitor is in a hurry, the server generates a response such as, "You have a package to deliver. If you are in a hurry, please scan the QR code."
[1958] Terminal: The generated response is transmitted to the visitor through a speaker, and a QR code is displayed on the digital signage.
[1959] QR code payment
[1960] User: Scans the displayed QR code using a smartphone and makes a payment.
[1961] Terminal: Once the payment is confirmed, the terminal will again announce to the visitor, "Payment has been confirmed. Please leave your luggage."
[1962] Recording and Monitoring
[1963] Terminal: Video and audio recording of interactions with visitors and sends the data to a cloud server.
[1964] Server: Analyzes stored data and notifies the user if any suspicious activity is detected.
[1965] Specific examples
[1966] Example: This system is installed on the door of a home where elderly people live.
[1967] Visitor detection and response
[1968] Device: A delivery person arrives at the elderly person's home. When the Edge AI camera detects the delivery person, the system is activated. A digital human appears and asks the delivery person, "Welcome. What can I do for you?"
[1969] User: The delivery person replies, "This is a courier delivery."
[1970] Speech recognition and intent understanding
[1971] Device: A microphone captures the voice and sends it to a cloud server. A generative AI model analyzes the voice and understands the intent of "delivery."
[1972] emotion recognition
[1973] Server: The emotion engine analyzes the delivery person's facial expressions and tone of voice to recognize emotions such as "I'm in a hurry."
[1974] Response generation and display
[1975] Server: The generative AI model generates a response saying, "You have a package delivery. If it's urgent, please scan the QR code." The speaker plays the message to the visitor. The digital sign displays the QR code.
[1976] QR code payment
[1977] User: The delivery person scans the QR code with their smartphone and makes the payment. After the payment is complete, the delivery person is told again by voice, "The payment has been confirmed. Please leave your package."
[1978] Recording and Monitoring
[1979] Device: Video and audio recordings during the call are saved on a cloud server. If any suspicious activity is detected, a notification is sent to the user's smartphone.
[1980] Example prompt sentence:
[1981] An example of a prompt sentence that a visitor can enter into the system using the API is, "This is a courier delivery. It's urgent, so please respond quickly."
[1982] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to safely and comfortably attend to visitors, and furthermore, it is possible to respond in a way that takes into consideration the feelings of visitors, providing a more user-friendly system.
[1983] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1984] Step 1: Visitor detection
[1985] Device: The Edge AI camera detects visitor movement. The camera constantly monitors the surroundings, and when movement is detected, the system automatically activates. Specifically, it identifies human silhouettes and movements and extracts their coordinate data.
[1986] Input: Camera video data
[1987] Output: coordinate data of detected motion and visitor detection trigger signal
[1988] Step 2: Voice capture and digital human response
[1989] Terminal: The microphone begins to capture the visitor's voice. At the same time, a digital human appears on the digital signage and asks, "Welcome. How can I help you?" The captured voice is converted into audio data through a signal processing engine.
[1990] Input: Visitor voice, visitor detection trigger signal
[1991] Output: Converted audio data, activation of digital human representation
[1992] Step 3: Visitor response
[1993] User: The visitor answers with their request, such as "This is a courier delivery." The microphone captures their voice.
[1994] Input: Visitor's response voice
[1995] Output: Captured response audio data
[1996] Step 4: Sending audio data
[1997] Device: The device sends the captured audio data to a cloud server, where it is converted into packets and transmitted using a secure protocol.
[1998] Input: Captured response audio data
[1999] Output: Send audio data to cloud server
[2000] Step 5: Voice data analysis and intent understanding
[2001] Server: The AI model generated on the cloud server analyzes the voice data and understands the visitor's intent. The voice data is converted into text data through a natural language processing engine, and the intent is extracted from the text data. For example, the intent "delivery" is understood from the phrase "This is a courier delivery."
[2002] Input: Audio data sent to the cloud server
[2003] Output: Text data about the visitor's intent (e.g., "delivery")
[2004] Step 6: Sending Emotion Data
[2005] Device: Video captured by the edge AI camera and audio captured by the microphone are sent to the emotion engine on the cloud server. The video and audio data are integrated and sent as a single data packet.
[2006] Input: Video and audio data
[2007] Output: Sending sentiment analysis data to the cloud server
[2008] Step 7: Sentiment Data Analysis
[2009] Server: The emotion engine recognizes emotions from the visitor's facial expressions and tone of voice. It extracts facial features from video data and analyzes tone characteristics from audio data. For example, it recognizes emotions such as "I'm in a hurry."
[2010] Input: Video and audio data sent to the cloud server
[2011] Output: Data about visitor sentiment (e.g., "I'm in a hurry")
[2012] Step 8: Response Generation
[2013] Server: Generates appropriate responses based on information obtained through speech analysis and emotion recognition. The generative AI model generates the optimal answer based on the visitor's intent and emotion data. For example, it generates a response such as, "You're delivering a package. If it's urgent, please scan the QR code."
[2014] Input: Visitor intent data, sentiment data
[2015] Output: Text data of the response
[2016] Step 9: Response display and communication
[2017] Terminal: The generated response is transmitted to the visitor via a speaker. A QR code is displayed on the digital signage. The response text is converted into speech using a speech synthesis engine and output from the speaker.
[2018] Input: Text data of response content
[2019] Output: Audio and visual response (QR code display)
[2020] Step 10: Scan the QR code and pay
[2021] User: Scans the displayed QR code using a smartphone and makes the payment. Once the payment is completed through the payment app, a notification is automatically sent to the system.
[2022] Input: QR code
[2023] Output: Payment completion notification
[2024] Step 11: Payment confirmation and instructions
[2025] Terminal: Confirm that the payment has been completed and tell the visitor again by voice, "Payment has been confirmed. Please leave your luggage."
[2026] Input: Payment completion notification
[2027] Output: Voice notification of payment completion
[2028] Step 12: Data recording
[2029] Device: Video and audio recordings are made during the call and sent to a cloud server. The recorded data is encrypted and stored securely.
[2030] Input: Video and audio data during the call
[2031] Output: Send recorded data to cloud server
[2032] Step 13: Detect and notify suspicious behavior
[2033] Server: Analyzes stored data and detects suspicious behavior. It uses algorithms to identify anomalous patterns of behavior and notifies the user if any are detected.
[2034] Input: Stored video and audio data
[2035] Output: Suspicious behavior detected notification
[2036] (Application example 2)
[2037] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2038] Currently, there are many systems that automatically respond to visitors, but few can recognize the visitor's emotions and respond appropriately. Furthermore, only a limited number of systems have the functionality to record the results of the response and detect suspicious behavior. Therefore, there is a demand for a system that can respond to visitors efficiently and safely, and that generates responses that take the visitor's emotions into consideration.
[2039] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for detecting when a visitor is approaching, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotions, means for outputting the generated response by voice and visually displaying it, means for recording the interaction with the visitor, and means for saving the recorded interaction and analyzing it as needed. This enables appropriate interaction that takes the visitor's emotions into consideration, thereby achieving safe and efficient visitor interaction.
[2040] "Means for detecting approaching visitors" refers to devices or software that detect the movement of visitors and activate the system.
[2041] "Means for capturing visitor audio" means a microphone or recording device used to collect and record audio emitted by a visitor.
[2042] "Means for analyzing captured voice data to understand the visitor's intent" refers to algorithms or software that analyzes voice data using voice recognition technology and understands the visitor's intent and requests.
[2043] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on collected data.
[2044] "Means for outputting the generated response audibly and displaying it visually" refers to equipment or software for outputting the generated response audibly through a speaker and displaying it on a display such as digital signage.
[2045] "Visitor interaction recording means" means any device or software that records visitor interactions in video or audio format.
[2046] "Means for storing the recorded response status and analyzing it as needed" refers to software and hardware for storing and analyzing the recorded data in a database or cloud server.
[2047] "Means for displaying QR codes" refers to devices or software that generate QR codes that encode specific information and display them on a display.
[2048] "Means for detecting and responding to suspicious behavior" refers to algorithms or software that analyze recorded data, detect abnormalities or suspicious behavior, and take appropriate action.
[2049] The present invention is a system for automatically responding to visitors, recognizing their emotions, and responding appropriately. This system includes means for detecting when a visitor approaches, means for capturing the visitor's voice, means for analyzing the captured voice data to understand the visitor's intention, means including a generative AI model for generating a response based on the understood intention and the visitor's emotion, means for outputting the generated response as voice and visually displaying it, means for recording the visitor interaction, and means for saving the recorded interaction and analyzing it as needed.
[2050] System Programming and Processing
[2051] The operation of this system is described as follows: This system uses the following hardware and software:
[2052] Hardware: Smartphone (camera, microphone), server, display, speaker
[2053] Software: OpenCV (image processing), SpeechRecognition (voice recognition), Transformers (generative AI model), TensorFlow (emotion recognition), qrcode (QR code generation), Firebase (data storage and analysis)
[2054] How to detect approaching visitors:
[2055] The system uses the smartphone camera to detect the movement of visitors. For example, when the smartphone camera recognizes a visitor, the system is activated.
[2056] To capture visitor audio:
[2057] The microphone on the smartphone is used to capture the visitor's voice. For example, if a visitor says, "I want to see new products," the microphone will collect that voice.
[2058] How to analyze captured audio data to understand visitor intent:
[2059] The collected voice data is converted into text data using the SpeechRecognition library. A generative AI model is then used to understand the visitor's intent. For example, a speech that says "I want to see new products" is understood to mean "I'm looking for product information."
[2060] Measures including generative AI models that generate responses based on the understood intent and sentiment of the visitor:
[2061] A response is generated based on the understood intent and the visitor's emotion (e.g., "excited") analyzed using an emotion recognition model using TensorFlow. Transformers is used as the generative AI model to input appropriate prompt sentences. For example, if the visitor is excited, the response generated will be, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[2062] A way to both speak and visually display the generated response:
[2063] The generated response is output as audio through the smartphone's speaker and displayed visually on the display, along with a QR code. For example, the QR code may be displayed along with a voice message saying, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[2064] How we record visitor interactions:
[2065] Visitor interactions are recorded in video and audio format, for example when a visitor scans a QR code.
[2066] A means to store recorded interactions and analyze them as needed:
[2067] The recorded data is stored in the cloud using Firebase and analyzed as needed to detect suspicious behavior. For example, if abnormal movements or behavior are detected, an administrator will be notified.
[2068] Examples and prompts
[2069] As a concrete example, consider a system placed at the entrance of a physical store. When a visitor approaches the store entrance, the system activates and analyzes the visitor's facial expressions and voice to respond. For example, along with a voice response such as "Welcome. How can I help you?", a response may be generated that instructs the visitor, "Here is a list of new products. Please choose your favorite product and scan the QR code."
[2070] Prompt Sentence Examples
[2071] "The visitor's sentiment is 'Excited' and they say: 'I want to see new products.' Generate an appropriate response. Visitor: 'I want to see new products.' Example of an appropriate response: 'Here is a list of new products. Choose your favorite and scan the QR code.'"
[2072] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2073] Step 1:
[2074] A means (device) of detecting approaching visitors
[2075] The system uses the smartphone camera to detect the movement of visitors. Specifically, it uses the OpenCV library to analyze the camera footage and perform motion detection. When motion is detected, the system is activated.
[2076] Input: Camera video data
[2077] Output: A signal that a visitor is approaching
[2078] Step 2:
[2079] A means (device) for capturing visitor audio
[2080] It uses the device's microphone to capture the visitor's voice, specifically by using the SpeechRecognition library to collect the voice data.
[2081] Input: Visitor utterance
[2082] Output: Audio data
[2083] Step 3:
[2084] A means of analyzing voice data and understanding the visitor's intent (server)
[2085] The captured voice data is sent to the server, where it is converted into text using the SpeechRecognition library, and then generative AI models (Transformers) are used to understand the visitor's intent.
[2086] Input: Audio data
[2087] Output: Text data containing visitor intent
[2088] Step 4:
[2089] A means of analyzing visitor sentiment (server)
[2090] Using video data captured on the server, we analyze visitors' emotions using a TensorFlow model, specifically identifying emotions from facial expressions and tone of voice.
[2091] Input: Video and audio data
[2092] Output: Visitor sentiment data
[2093] Step 5:
[2094] A means (server) to generate responses based on the understood intent and emotions
[2095] Using generative AI models (Transformers), it generates appropriate responses based on the visitor's intent and emotions. For example, if the visitor is excited, it might generate a response like, "Here's a list of new products. Choose your favorite and scan the QR code."
[2096] Input: Text data containing intent, emotion data
[2097] Output: The generated response text
[2098] Step 6:
[2099] A means (terminal) to output the generated response as voice and display it visually
[2100] The generated response text is synthesized into speech and output as voice through the smartphone speaker. The response text including a QR code is also displayed on the display. The QR code is generated using the qrcode library.
[2101] Input: Generated response text
[2102] Output: Audio output, visual display with QR code
[2103] Step 7:
[2104] A means (terminal) for recording visitor interactions
[2105] Record visitor interactions in video and audio format using the smartphone camera and microphone.
[2106] Input: Video and audio when greeting a visitor
[2107] Output: Recorded video and audio data
[2108] Step 8:
[2109] A means (server) to store the recorded response status and analyze it as needed
[2110] The recorded video and audio data is stored on the Firebase cloud server. If necessary, the stored data is analyzed, and if any suspicious activity is detected, an administrator is notified.
[2111] Input: Recorded video and audio data
[2112] Output: Data stored in the cloud, suspicious behavior detection results
[2113] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2114] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2115] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2116] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2117] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2118] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2119] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2120] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2121] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2122] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2123] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2124] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2125] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2126] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2127] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2128] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2129] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2130] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2131] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2132] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2133] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2134] The following is further disclosed regarding the above embodiment.
[2135] (Claim 1)
[2136] a means for detecting approaching visitors;
[2137] a means for capturing the visitor's voice;
[2138] A means of analyzing voice data to understand visitor intent,
[2139] a means for generating a response based on the understood intent;
[2140] means for audibly outputting the generated response;
[2141] A means for recording visitor interactions;
[2142] A means for storing the recorded response status;
[2143] A system including:
[2144] (Claim 2)
[2145] 10. The system of claim 1, wherein the response generating means further comprises: means for displaying a QR code after understanding the visitor's intent.
[2146] (Claim 3)
[2147] 10. The system of claim 1, further comprising means for detecting suspicious behavior based on the recorded interactions.
[2148] "Example 1"
[2149] (Claim 1)
[2150] a means for detecting approaching visitors;
[2151] a means for capturing the visitor's voice;
[2152] A means of analyzing voice data to understand visitor intent,
[2153] a means for generating a response based on the understood intent;
[2154] means for audibly outputting the generated response;
[2155] a means of displaying visual information to visitors;
[2156] A means for recording visitor interactions;
[2157] A means for storing the recorded response status;
[2158] A system including:
[2159] (Claim 2)
[2160] 10. The system of claim 1, wherein the response generating means further comprises: means for displaying a QR code after understanding the visitor's intent.
[2161] (Claim 3)
[2162] 10. The system of claim 1, further comprising means for detecting suspicious behavior based on the recorded interactions.
[2163] "Application Example 1"
[2164] (Claim 1)
[2165] a means for detecting approaching visitors;
[2166] a means for capturing the visitor's voice;
[2167] A means of analyzing voice data to understand visitor intent,
[2168] a means for generating a response based on the understood intent;
[2169] means for audibly outputting the generated response;
[2170] A means for recording visitor interactions;
[2171] A means for storing the recorded response status;
[2172] a means for employees equipped with smart glasses to visually confirm the response; and
[2173] A system including:
[2174] (Claim 2)
[2175] 10. The system of claim 1, wherein the response generating means further comprises: means for displaying a QR code after understanding the visitor's intent.
[2176] (Claim 3)
[2177] 10. The system of claim 1, further comprising means for detecting suspicious behavior based on the recorded interactions.
[2178] "Example 2: Combining Emotion Engines"
[2179] (Claim 1)
[2180] a means for detecting approaching visitors;
[2181] a means for capturing the visitor's voice;
[2182] means for transmitting the voice data to a cloud server;
[2183] A means to analyze voice data on a cloud server and understand the visitor's intent,
[2184] a means for performing emotion recognition based on the understood intent;
[2185] means for generating a response based on the emotion recognition result and the intention;
[2186] means for visually and audibly outputting the generated response;
[2187] a means for displaying a QR code as part of the response;
[2188] A means for recording visitor interactions;
[2189] A means for storing the recorded response status in a cloud server;
[2190] A system including:
[2191] (Claim 2)
[2192] 10. The system of claim 1, wherein the response generating means further comprises means for selecting optimal response content based on the visitor's intent and emotion recognition results.
[2193] (Claim 3)
[2194] The system according to claim 1, further comprising means for analyzing and detecting suspicious behavior on a cloud server based on the recorded response situations.
[2195] "Application example 2 when combining emotion engines"
[2196] (Claim 1)
[2197] a means for detecting approaching visitors;
[2198] a means for capturing the visitor's voice;
[2199] A means of analyzing captured voice data to understand visitor intent,
[2200] a generative AI model that generates a response based on the understood intent and sentiment of the visitor;
[2201] means for audibly outputting and visually displaying the generated response;
[2202] A means for recording visitor interactions;
[2203] A means of storing the recorded response status and analyzing it as needed;
[2204] A system including:
[2205] (Claim 2)
[2206] 10. The system of claim 1, further comprising: means for displaying the QR code after generating an appropriate response based on the understood intent and emotion.
[2207] (Claim 3)
[2208] 10. The system of claim 1, further comprising means for detecting and responding to suspicious activity based on the recorded interactions. [Explanation of symbols]
[2209] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for detecting approaching visitors; a means for capturing the visitor's voice; A means of analyzing voice data to understand visitor intent, a means for generating a response based on the understood intent; means for audibly outputting the generated response; A means for recording visitor interactions; A means for storing the recorded response status; A system including:
2. The system of claim 1 , wherein the response generating means further comprises: means for displaying a QR code after understanding the visitor's intent.
3. The system of claim 1 , further comprising means for detecting suspicious behavior based on the recorded responses.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A