system
The digital human reception system addresses visitor interaction challenges by automating attribute-based responses and secure record management, enhancing safety and reducing user effort.
Patent Information
- Application Number
- JP2024123935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
The increase in telecommuting and online shopping has led to more interactions with visitors, posing risks and discomfort for elderly, disabled, and children, especially when dealing with solicitors or delivery persons, as existing systems struggle with accurate visitor attribute identification, flexible interactions, and secure record management.
A digital human reception system that captures visitor video, analyzes attributes, generates digital human interactions, records conversations, and provides QR code processing, using AI and camera technology to automate and secure interactions.
The system reduces user workload and risk by automating visitor interactions, providing appropriate responses based on attributes, and ensuring secure record keeping.
Smart Images

Figure 2026022418000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's society, the spread of telecommuting and online shopping has led to an increase in interactions with visitors through doors. Furthermore, dealing with visitors directly often poses risks for the elderly, people with disabilities, and children. When dealing with visitors when the user is not at home, or when the visitor is a solicitor or delivery person, there is a risk of trouble occurring and the user feels uncomfortable. To solve these problems, a new system is needed that reduces the hassle and risk for users and enables safe and comfortable interactions with visitors. [Means for solving the problem]
[0005] The present invention provides a digital human reception system that automates interactions with visitors and reduces the user's workload and risk. The system of the present invention includes the following means.
[0006] The system includes a means for capturing video of nearby visitors using a camera device, a means for analyzing the video data and identifying the visitor's attributes, a means for generating video and audio of a digital human based on the attributes, a means for interacting with the visitor, and a means for recording the content of the interaction and the visitor's video. The system also includes a means for generating and displaying a QR code.
[0007] A system configured in this way automates interactions with visitors, reducing the effort and risk for users and enabling appropriate responses based on the visitor's attributes.
[0008] A "camera device" is an electronic device for capturing an image of an object and transmitting that data to another device.
[0009] "Video data" refers to the digital information of images and videos captured by a camera device.
[0010] "Attributes" are information that indicates the visitor's external characteristics (e.g., age, gender, belongings, etc.).
[0011] A "digital human" is an artificially generated image and voice of a human being that is used to interact with visitors.
[0012] "Video analysis" is a process for extracting necessary information from video data and detecting specific attributes.
[0013] "Speech synthesis" is a technology that generates natural-sounding speech from text data.
[0014] An "audio recording device" is a device for capturing ambient sounds and storing or transmitting them as digital data.
[0015] A "voice recognition module" is technology or software that analyzes voice data and converts it into text.
[0016] A "QR code" is a type of two-dimensional barcode that has a pattern that allows information to be read efficiently.
[0017] "Dialogue content" refers to the content of the conversation between the visitor and the digital human.
[0018] A "database" is a system for systematically storing and managing digital data. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[0041] System Configuration
[0042] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[0043] 2. Server: Receives and analyzes the video data sent from the camera device, identifies visitor attributes, and generates video and audio of a digital human based on those attributes.
[0044] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[0045] 4. User Interface: This is the interface that allows users to configure the system and check records.
[0046] Program processing explanation
[0047] 1. Visitor Recognition
[0048] The terminal transmits the video data from the camera device to the server.
[0049] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0050] 2. Digital Human Generation
[0051] The server generates the appropriate video and audio of the digital human based on the identified attributes.
[0052] For example, the setting is such that elderly people speak in a slow tone, while children speak in a friendly tone.
[0053] 3. Interacting with visitors
[0054] The terminal displays the digital human received from the server and greets the visitor.
[0055] The visitor speaks to the digital human.
[0056] The terminal picks up the visitor's voice and transmits it to the server.
[0057] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[0058] The terminal communicates the generated answer to the visitor via a digital human.
[0059] 4. QR Code Processing
[0060] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[0061] The server generates a QR code and sends it to the device.
[0062] The terminal displays a QR code that visitors can scan with their smartphone.
[0063] 5. Recording and Confirmation
[0064] The server records all conversations and videos.
[0065] Users can later review the recorded conversations and footage through a dedicated interface.
[0066] Specific examples
[0067] Visitor recognition and digital human generation
[0068] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[0069] Interacting with visitors
[0070] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates the appropriate response, "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[0071] QR code processing and recording
[0072] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[0073] The above is an embodiment of the present invention. The present invention automates the process of dealing with visitors, significantly reducing the effort and risk for users.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[0077] Step 2:
[0078] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[0079] Step 3:
[0080] The server generates the appropriate video and audio of a digital human based on the identified attributes, for example, an elderly person will have a slower voice.
[0081] Step 4:
[0082] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[0083] Step 5:
[0084] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[0085] Step 6:
[0086] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[0087] Step 7:
[0088] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[0089] Step 8:
[0090] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[0091] Step 9:
[0092] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[0093] Step 10:
[0094] The visitor tells the digital human that they would like to process the transaction using a QR code.
[0095] Step 11:
[0096] The server generates a QR code and sends it to the device.
[0097] Step 12:
[0098] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[0099] Step 13:
[0100] The server records all interactions and videos with visitors and stores them in a database.
[0101] Step 14:
[0102] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[0103] Example 1
[0104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0105] There is a need to build a system that automates visitor interactions and provides appropriate responses based on the visitor's attributes. Conventional visitor interaction systems have difficulty accurately identifying attributes such as the visitor's age and gender and generating a suitable digital human. Furthermore, many systems lack the ability to record and review the content and video of interactions with visitors. Furthermore, there is a lack of a convenient way to make payments and receive items using QR codes.
[0106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0107] In this invention, the server includes means for analyzing video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, and means for generating a greeting for the digital human using a generative AI model. This makes it possible to generate and display an appropriate digital human according to the visitor attributes, automatically generate a greeting and response according to the attributes, and record and check the content of conversations with visitors and their video.
[0108] A "camera device" is a photographic device for capturing images of visitors, and is installed outside the door, etc.
[0109] "Video data" refers to video information of visitors captured by a camera device.
[0110] "Analyzing video data" means performing processing to identify visitor attribute information (age, gender, belongings, etc.) from the received video data.
[0111] "Visitor attributes" refers to characteristic information about a visitor, such as age, gender, and belongings.
[0112] "Digital Human" means a virtual character, including video and audio, that is generated to interact with visitors.
[0113] An "audio recording device" is a device such as a microphone for capturing the visitor's voice.
[0114] "Voice data" refers to audio information that records the speech of a visitor.
[0115] "Speech recognition technology" is a technology for converting voice data into text and analyzing the content of speech.
[0116] A "generative AI model" is an artificial intelligence model that generates appropriate responses and greetings based on attribute information and speech content.
[0117] A "QR code" is a two-dimensional barcode that visitors use when making payments or receiving items.
[0118] "Recording" refers to the act of saving the content of conversations with visitors and footage.
[0119] "User interface" refers to the interface that allows the user to configure the system and check the recorded dialogue and video.
[0120] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[0121] System Configuration
[0122] 1. Camera device: Installed on the exterior of the door to capture video of visitors. It has the ability to capture high-resolution video, capturing detailed images of visitors' faces and belongings.
[0123] 2. Server: Receives and analyzes the video data sent from the camera. The server incorporates image analysis software (e.g., OpenCV or TensorFlow) and speech recognition technology (e.g., Google Speech-to-Text API). It also generates digital human responses using generative AI models (e.g., OpenAI's GPT-3.5).
[0124] 3. Terminal: Receives data from the server and interacts with visitors. The terminal is equipped with a voice recording device that captures the visitor's voice and sends it to the server.
[0125] 4. User Interface: This is the interface through which users can configure the system, view recorded conversations, and check video footage. It is built as a web application and can be accessed from a browser (e.g., using React or Angular).
[0126] Program processing explanation
[0127] 1. Visitor Recognition:
[0128] The device sends the video data from the camera to the server, which analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0129] 2. Digital Human Generation:
[0130] The server then generates the appropriate video and audio of a digital human based on the identified attributes, and the generated digital human is configured to speak in different tones depending on the visitor's attributes, for example, speaking in a slower tone for elderly people and in a more friendly tone for children.
[0131] 3. Visitor interaction:
[0132] The device displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the device's audio pickup device and sent to the server. The server uses speech recognition technology to analyze the speech and generates an appropriate response using a generative AI model. The device then conveys this response to the visitor via the digital human.
[0133] 4. QR Code Processing:
[0134] When a visitor wishes to make or receive payment using a QR code, they tell the digital human, and the server generates a QR code and sends it to the terminal, which displays the code and the visitor scans it with their smartphone.
[0135] 5. Record and verify:
[0136] The server records all conversations and videos, and users can later review the recorded conversations and videos through a dedicated interface.
[0137] Specific examples
[0138] Visitor recognition and digital human generation
[0139] One day, an elderly visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model designed for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[0140] Interacting with visitors
[0141] The visitor says, "You have a delivery," and the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and generates an appropriate response based on the text: "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[0142] QR code processing and recording
[0143] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[0144] Prompt Sentence Examples
[0145] Here are some example prompts to input to a generative AI model (e.g., OpenAI's GPT-3.5):
[0146] Generate digital human greetings based on visitor attributes. Digital human greetings based on the following attributes:
[0147] Age: Elderly
[0148] Gender: Female
[0149] Possession: Cane
[0150] Please speak in a relaxed, friendly tone for seniors.
[0151] Using this prompt, the system is capable of generating an appropriate digital human greeting for the visitor.
[0152] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0153] Step 1:
[0154] Video capture using camera equipment
[0155] The terminal acquires video data from a camera device installed outside the door.
[0156] Input: Real-time video of the visitor captured by the camera device.
[0157] Output: Captured video data.
[0158] Step 2:
[0159] Video data sent to server
[0160] The terminal compresses the acquired video data, minimizes the amount of data, and transmits it to the server.
[0161] Input: Video data acquired from a camera device.
[0162] Output: Compressed video data sent to the server.
[0163] Step 3:
[0164] Identifying visitor attributes
[0165] The server uses image analysis software (e.g., OpenCV or TensorFlow) to analyze the received video data.
[0166] Through video analysis, the server identifies visitors' attributes such as age, gender, and belongings.
[0167] Input: Compressed video data.
[0168] Output: Visitor demographic information (e.g. age, gender, belongings).
[0169] Step 4:
[0170] Digital Human Generation
[0171] The server uses a generative AI model (e.g., OpenAI's GPT-3.5) to generate a greeting for the digital human based on the visitor's identified attribute information.
[0172] Next, 3D modeling software (e.g., Blender or Unity) is used to generate the video and audio of the digital human.
[0173] Input: Visitor demographic information.
[0174] Output: Digital human video and audio data.
[0175] Step 5:
[0176] Digital human display and greeting
[0177] The terminal displays the video and audio of the digital human received from the server.
[0178] A digital human greets visitors.
[0179] Input: Video and audio data of a digital human.
[0180] Output: A displayed digital human greeting.
[0181] Step 6:
[0182] Interacting with visitors
[0183] The visitor speaks to the digital human.
[0184] The terminal picks up the visitor's speech and sends the audio data to the server.
[0185] Input: Visitor utterance.
[0186] Output: The audio data sent to the server.
[0187] Step 7:
[0188] Voice data analysis and response generation
[0189] The server uses voice recognition technology (e.g., Google Speech-to-Text API) to convert the visitor's speech into text.
[0190] It then uses a generative AI model to generate appropriate responses to what the visitor says.
[0191] Input: Visitor's voice data.
[0192] Output: The recognized text and the response message.
[0193] Step 8:
[0194] Viewing the response
[0195] The terminal conveys the response message received from the server to the visitor through the digital human.
[0196] Input: Digital human's response message.
[0197] Output: Displayed digital human response.
[0198] Step 9:
[0199] QR code processing
[0200] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[0201] The server generates a dedicated QR code and sends it to the device.
[0202] The terminal displays a QR code that visitors can scan with their smartphone.
[0203] Input: Visitor processing request.
[0204] Output: The displayed QR code.
[0205] Step 10:
[0206] Dialogue and video recording
[0207] The server records all conversations and videos.
[0208] Input: Dialogue content and video data.
[0209] Output: Recorded dialogue and video data.
[0210] Step 11:
[0211] Checking the recorded content
[0212] Users can view the recorded conversations and footage through a dedicated user interface.
[0213] Input: Request to access recorded data.
[0214] Output: Displayed recording and video data.
[0215] (Application example 1)
[0216] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0217] Conventional visitor response systems have had several limitations in terms of visitor recognition, interaction, and enhanced security. Specifically, they have difficulty in flexibly interacting with visitors and responding quickly, which increases the workload of users. Furthermore, they have been unable to provide appropriate responses based on visitor attributes or automatically manage records. This has led to problems with systems not functioning effectively, particularly when individualized attention is required for elderly people and children.
[0218] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0219] In this invention, the server includes means for capturing video of nearby visitors with a camera device, means for analyzing the video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, means for interacting with the visitor, means for recording the content of the interaction and the visitor video, means for interacting with the visitor using a smartphone, means for capturing the visitor's audio with the smartphone and analyzing the audio data to generate an appropriate response, and means for generating and displaying a QR code on the smartphone. This enables flexible visitor recognition and appropriate response, reduces user effort, and enhances security.
[0220] A "camera device" is a device for capturing video of visitors in real time.
[0221] "Video data" refers to video information of visitors captured by a camera device.
[0222] "Visitor attributes" refers to information such as the visitor's age, gender, belongings, etc.
[0223] A "digital human" is a video and audio representation of a virtual person generated based on the visitor's attributes.
[0224] "Dialogue" refers to communication between a digital human and a visitor.
[0225] "Recording" refers to saving the content of conversations and footage of visitors.
[0226] A "smartphone" is a portable information terminal that works in conjunction with a camera device and a server and is used to interact with visitors.
[0227] "Audio data" refers to information that records the content of a visitor's speech.
[0228] A "QR code" is a two-dimensional barcode generated for visitors to make payments and other transactions.
[0229] The "server" is a central processing unit that receives data from the camera device, analyzes it, generates digital humans, and records them.
[0230] MODE FOR CARRYING OUT THE INVENTION
[0231] A system for implementing the present invention is configured as follows: First, a camera device captures video of nearby visitors in real time and transmits the video data to a server. Next, the server analyzes the video data and identifies the visitor's attributes. Based on the identified attributes, the server generates video and audio of a digital human and interacts with the visitor via a smartphone.
[0232] Required Hardware and Software
[0233] Camera equipment: Equipment for capturing images of visitors. High-resolution cameras are recommended.
[0234] Server: A central processing unit that analyzes data, generates digital humans, and records data. It incorporates machine learning libraries such as TensorFlow and PyTorch.
[0235] Smartphone: A mobile information device for interacting with visitors. It requires an internet connection, a camera, and a microphone.
[0236] Speech recognition and speech synthesis technologies: The server incorporates speech recognition technologies (such as the Google Speech-to-Text API) and speech synthesis technologies (such as the Google Text-to-Speech API).
[0237] QR code generation software: A library for generating QR codes (e.g., Python's QRCode library).
[0238] Processing flow
[0239] Step 1: Video capture
[0240] The camera captures video of visitors in real time and sends the video data to a server, which analyzes the received video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0241] Step 2: Creating a digital human
[0242] The server generates the appropriate video and audio of the digital human based on the identified attributes. For example, it may be configured to speak in a slower tone for an elderly person, or in a more friendly tone for a child. The data of the generated digital human is then sent to a smartphone.
[0243] Step 3: Dialogue with visitors
[0244] The smartphone displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the smartphone's microphone and sent to the server. The server uses voice recognition technology to analyze the speech and generate an appropriate response. The smartphone then broadcasts the generated response via the digital human.
[0245] Step 4: Processing by QR code
[0246] When a visitor wishes to make a payment or receive payment using a QR code, they notify the digital human. The server generates a dedicated QR code and sends it to the smartphone. The smartphone displays this QR code, and the visitor can scan it to complete the payment process.
[0247] Step 5: Record and verify
[0248] The server records all conversations and video footage, which users can later review through a dedicated interface. This automates interactions with visitors, significantly reducing the user's workload and risk.
[0249] Examples and prompts
[0250] Specific examples
[0251] One day, a visitor approaches the door of a home. A camera installed on the door captures the visitor's video and sends it to a server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people, generates video and audio of the visitor speaking in a slow tone, saying, "Hello. How can I help you?" and sends this video and audio to the smartphone. The smartphone then displays this digital human and greets the visitor.
[0252] Prompt Sentence Examples
[0253] "Generate the appropriate digital human based on the visitor's attributes. For example, set a slower tone for seniors and a more friendly tone for children. Also, generate a QR code."
[0254] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0255] Step 1:
[0256] Input: Video data captured by a camera device
[0257] Processing: The camera device captures video of nearby visitors in real time and collects the video data.
[0258] Output: Collected video data
[0259] Specific operation: The camera captures the moment a visitor approaches the door and saves the video data.
[0260] Step 2:
[0261] Input: Video data obtained from a camera device
[0262] Processing: The server receives the video data and uses video analytics techniques (e.g., facial and object recognition) to identify visitor attributes (e.g., age, gender, belongings, etc.).
[0263] Output: Visitor attribute data identified by analysis
[0264] Specific operation: The server analyzes the video data using TensorFlow and PyTorch, identifies the visitor's face and features, and determines their age and gender.
[0265] Step 3:
[0266] Input: Visitor attribute data
[0267] Processing: The server generates the appropriate video and audio of the digital human based on the visitor's attributes.
[0268] Output: Video and audio data of the generated digital human
[0269] How it works: The server uses the generative AI model to generate a digital human that, for example, greets elderly people in a slower tone. Video and audio data are generated.
[0270] Step 4:
[0271] Input: Digital human video and audio data
[0272] Processing: The server sends the generated digital human data to the smartphone, which displays the digital human to the visitor and greets them.
[0273] Output: Video and audio of a digital human displayed on a smartphone
[0274] Specific operation: The smartphone plays the video of the digital human it receives and greets the user with a greeting such as "Hello. How can I help you?"
[0275] Step 5:
[0276] Input: What the visitor said
[0277] Processing: The device picks up what the visitor is saying and sends the audio data to the server, which uses speech recognition technology to convert it into text data and generate an appropriate response.
[0278] Output: Speech-to-text data and the digital human's response based on it
[0279] Specific operation: The smartphone uses a microphone to pick up the visitor's speech, such as "I have a delivery for you," and the server converts this into text data and generates a response such as "Okay, you can pick it up using the QR code."
[0280] Step 6:
[0281] Input: Visitor's request (e.g., request for pickup via QR code)
[0282] Process: The server generates a QR code and sends it to the smartphone, which displays the generated QR code to the visitor.
[0283] Output: QR code displayed on the smartphone
[0284] How it works: The server generates a QR code using a Python QRCode library or similar, and the smartphone displays the QR code to the visitor.
[0285] Step 7:
[0286] Input: Dialogue content and video data
[0287] Processing: The server records all conversations and video. The user can view the recorded data through a dedicated interface.
[0288] Output: Recorded dialogue and video data
[0289] How it works: The server stores all the interaction data and footage in a database, allowing users to view this data later.
[0290] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0291] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, an emotion engine, and a user interface.
[0292] System Configuration
[0293] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[0294] 2. Server: Receives and analyzes the video and audio data sent from the camera device, identifies the visitor's attributes and emotions, and generates the video and audio of a digital human based on those.
[0295] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[0296] 4. Emotion Engine: An engine that analyzes the visitor's video and audio data and recognizes the visitor's emotions.
[0297] 5. User Interface: This is the interface that allows users to configure the system and check records.
[0298] Program processing explanation
[0299] 1. Visitor Recognition
[0300] The terminal transmits the video data from the camera device to the server.
[0301] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0302] 2. Emotional Recognition
[0303] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[0304] Identify different emotional states of your visitors, such as happy, angry, or anxious.
[0305] 3. Digital Human Generation
[0306] The server generates the appropriate video and audio of the digital human based on the identified attributes and emotions.
[0307] For example, if an elderly visitor is feeling anxious, you might adopt a more friendly and gentle tone.
[0308] 4. Interacting with visitors
[0309] The terminal displays the digital human received from the server and makes initial responses to visitors, such as greetings.
[0310] Visitors speak to the digital human.
[0311] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[0312] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[0313] 5. Regulating responses based on emotions
[0314] Based on the analysis results of the emotion engine, the server adjusts the visual and audio tone of the digital human, for example, responding in a calmer tone to soothe the visitor's anger.
[0315] 6. QR Code Processing
[0316] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[0317] The server generates a QR code and sends it to the device.
[0318] The terminal displays a QR code that visitors can scan with their smartphone.
[0319] 7. Recording and Confirmation
[0320] The server records all conversations and videos.
[0321] The recording also includes visitor sentiment data, which users can review later.
[0322] Users can later review the recorded conversations and footage through a dedicated interface.
[0323] Specific examples
[0324] Visitor recognition and emotion recognition
[0325] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. It also uses an emotion engine to identify that the visitor is feeling anxious. The server then generates a digital human with a gentle tone of voice, specifically for the elderly, to ease their anxiety.
[0326] Interact with visitors and tailor responses based on their emotions
[0327] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates an appropriate response, "Okay, you can pick it up by scanning the QR code." Based on the analysis results of the emotion engine, the response is adjusted to a gentler tone. The device then plays this response on a digital human and conveys it to the visitor.
[0328] QR code processing and recording
[0329] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the device. The device displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server and saved, including emotional data. Users can later review these records through a dedicated interface.
[0330] The above is an embodiment of the present invention. This invention automates the process of interacting with visitors, significantly reducing the user's workload and risk. By combining it with an emotion engine, it becomes possible to respond appropriately to visitors' emotional states.
[0331] The processing flow will be explained below.
[0332] Step 1:
[0333] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[0334] Step 2:
[0335] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[0336] Step 3:
[0337] The server analyzes the visitor's emotions from the video and audio data using an emotion engine, which identifies the visitor's emotional state from their facial expressions and tone of voice.
[0338] Step 4:
[0339] The server generates an appropriate digital human image and voice based on the identified attributes and emotions. For example, a digital human with a gentle tone is generated for an elderly visitor who is anxious.
[0340] Step 5:
[0341] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[0342] Step 6:
[0343] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[0344] Step 7:
[0345] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[0346] Step 8:
[0347] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[0348] Step 9:
[0349] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[0350] Step 10:
[0351] Based on the analysis results of the emotion engine, the server adjusts the video and audio tones of the digital human and responds according to the visitor's emotional state.
[0352] Step 11:
[0353] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[0354] Step 12:
[0355] The visitor tells the digital human that they would like to process the transaction using a QR code.
[0356] Step 13:
[0357] The server generates a QR code and sends it to the device.
[0358] Step 14:
[0359] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[0360] Step 15:
[0361] The server records all interactions and videos with visitors and stores them in a database, including emotional data analyzed by the emotion engine.
[0362] Step 16:
[0363] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[0364] Example 2
[0365] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0366] Automated response systems that interact with visitors are required to accurately recognize the visitor's attributes and emotions and generate an appropriate digital human. It is also important to enable visitors to smoothly make payments and receive payments using QR codes. However, integrating these elements into a single system and making it function smoothly is technically difficult.
[0367] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0368] In this invention, the server includes means for capturing video of nearby visitors using a camera device, means for analyzing the video data and identifying the visitor's attributes, means for analyzing the visitor's emotions from the video and audio data, means for generating video and audio of a digital human based on the visitor's attributes and emotions, means for displaying the digital human and engaging in a dialogue with the visitor, means for capturing the visitor's audio and generating an appropriate response using voice recognition technology, means for adjusting the digital human's response based on the emotion analysis results, means for generating and displaying a QR code when the visitor uses a QR code to make or receive payments, and means for recording the content of the dialogue and the visitor's video. This allows the visitor's attributes and emotions to be accurately recognized, enabling smooth payment and receipt using an appropriate digital human for dialogue and the QR code.
[0369] "Camera Device" means equipment used to capture real-time video of visitors.
[0370] "Video Data" refers to still and video images of Visitors captured by camera devices.
[0371] The "server" is a central processing unit that analyzes video and audio data and identifies the attributes and emotions of visitors.
[0372] "Visitor attributes" refers to characteristic information about visitors, such as age, gender, and belongings.
[0373] The "Emotion Engine" is a software interface for analyzing a visitor's emotional state from video and audio data.
[0374] A "digital human" is an artificial visual and audio character that is generated to interact with visitors.
[0375] "Voice recognition technology" is a technology for converting a visitor's voice into text data.
[0376] A "QR code" is a two-dimensional barcode used by visitors to make payments and receive goods.
[0377] "Dialogue content" refers to the entire content of the conversation between the visitor and the digital human.
[0378] "Recording means" refers to methods and devices for saving the content of interactions and video of visitors.
[0379] In one embodiment of the present invention, the system is composed of a camera device, a server, a terminal, an emotion engine, and a user interface, which allows the system to recognize the attributes and emotions of visitors and generate an appropriate digital human to interact with them.
[0380] System Configuration
[0381] camera equipment
[0382] The camera device is installed outside the door and captures visitors' images in real time. The camera device acquires high-resolution images and transmits them to a terminal.
[0383] server
[0384] The server receives and analyzes the video and audio data sent from the camera. The server has the following functions:
[0385] 1. Video analysis: Using a video analysis library such as OpenCV, we identify visitors' faces and extract attributes such as age, gender, and belongings.
[0386] 2. Emotion analysis: Using Microsoft Azure's emotion recognition API, we analyze visitors' emotions from video and audio data.
[0387] 3. Digital Human Generation: Using tools such as Adobe Character Animator, generate the video and audio of a digital human based on identified attributes and emotions.
[0388] 4. Speech Recognition: Using APIs such as Google Cloud Speech-to-Text, convert the visitor's voice data into text and generate an appropriate response.
[0389] 5. Response adjustment: Adjust the digital human's response based on the emotion analysis results.
[0390] Terminal
[0391] The terminal is the device that receives data from the server and interacts with the visitor. The terminal has the following functions:
[0392] 1. Video display: Displays the digital human received from the server and provides initial greetings and guidance to visitors.
[0393] 2. Microphone: A built-in microphone is included to capture the visitor's speech and send it to the server.
[0394] 3. Display QR code: When a visitor uses a QR code to make or receive a payment, the QR code sent from the server is displayed.
[0395] Emotion Engine
[0396] The emotion engine analyzes the video and audio data of visitors and recognizes their emotions. This engine can identify their emotional state, such as whether they are happy, angry, or anxious.
[0397] User Interface
[0398] Users can use the interface to configure the system and view recordings, which include visitor attributes, emotional state, dialogue, and video.
[0399] Specific operation example
[0400] 1. Recognizing visitors' perceptions and emotions
[0401] One day, a visitor approaches the door of your home. A camera installed on the door captures the visitor's image, and the device sends it to the server.
[0402] The server uses video analysis to determine that the visitor is elderly, and an emotion engine to identify that the visitor is feeling anxious.
[0403] The server is designed for the elderly, generating a digital human with a gentle tone to ease anxiety.
[0404] 2. Interact with visitors and tailor responses based on their emotions
[0405] When a visitor says "I have a delivery," the voice is captured by the device's microphone and sent to the server.
[0406] The server uses speech recognition technology to generate the text "You have a delivery" and then generates the appropriate response "Okay, you can pick it up using the QR code."
[0407] Based on the analysis results of the emotion engine, the response is adjusted to a gentle tone, which the device then plays back to the digital human and conveys to the visitor.
[0408] 3. Processing and recording by QR code
[0409] When a visitor requests to receive their gift using a QR code, the server generates a dedicated QR code and sends it to the device.
[0410] The terminal displays this QR code, and visitors can scan it with their smartphone to complete the pickup.
[0411] All conversations and videos are recorded on a server, including emotional data, and users can view these records through a dedicated interface.
[0412] Prompt Sentence Examples
[0413] 1. Prompt to identify visitor demographics:
[0414] Identify the visitor's age, gender, and belongings from the video data, for example, whether the visitor is carrying a bag.
[0415] 2. Emotion analysis prompts:
[0416] Analyze visitor emotions from this video and audio data. Identify whether your visitors are happy, angry, anxious, etc.
[0417] 3. Prompt when generating a digital human:
[0418] If your visitor is elderly and anxious, generate a digital human that responds in a friendly and gentle tone.
[0419] 4. Speech recognition prompts:
[0420] Convert the visitor's speech into text data. For example, convert the speech "I have a delivery" into text.
[0421] The above is an embodiment of the present invention. The present invention enables accurate recognition of visitor attributes and emotions, and enables smooth payment and receipt using appropriate digital human dialogue and QR codes.
[0422] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0423] Step 1:
[0424] The terminal acquires video data from the camera device and transmits it to the server.
[0425] Input: Real-time video data from a camera device
[0426] Data processing: The device reads the video data from the buffer.
[0427] Output: Video data transmitted to the server through the configured network connection
[0428] Specific operation: A camera device is installed outside the door, and when a visitor approaches, it automatically captures video and the device sends the video data to the server.
[0429] Step 2:
[0430] The server analyzes the received video data and identifies the visitor's attributes.
[0431] Input: Video data sent from the device
[0432] Data processing: Using a video analysis library such as OpenCV, a facial recognition algorithm is applied to extract attributes such as age, gender, and belongings.
[0433] Output: Visitor attribute data (e.g., age: 60, gender: male, belongings: bag)
[0434] What happens: The server performs video analysis to detect the visitor's face and identify their attributes.
[0435] Step 3:
[0436] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[0437] Input: Analyzed video and audio data
[0438] Data processing: Analyze facial expressions and tone of voice using Microsoft Azure's emotion recognition API.
[0439] Output: Visitor emotion data (e.g., anxiety, joy, anger)
[0440] What happens: The server sends data to the emotion engine to identify the visitor's emotion.
[0441] Step 4:
[0442] The server generates a digital human based on the visitor's identified attributes and emotions.
[0443] Input: Visitor demographic and sentiment data
[0444] Data processing: Using Adobe Character Animator or similar software, we generate appropriate digital human images and voices.
[0445] Output: Video and audio data of the generated digital human
[0446] Specific behavior: The server generates a digital human with a friendly and gentle tone that corresponds to the identified attributes (e.g., elderly, anxious state).
[0447] Step 5:
[0448] The terminal displays the digital human received from the server and makes an initial response to the visitor.
[0449] Input: Video and audio data of the digital human sent from the server
[0450] Data processing: Display and play data received by the terminal
[0451] Output: The video and audio of the digital human are displayed and played back to the visitor.
[0452] What it does: The device greets the visitor with the video and audio of a digital human, asking, "Hello, how can I help you?"
[0453] Step 6:
[0454] The visitor speaks to the digital human, and the device captures the audio and sends it to the server.
[0455] Input: Visitor's speech
[0456] Data processing: The device's microphone captures the audio and sends the audio data to the server.
[0457] Output: Audio data sent to the server
[0458] Specific operation: The visitor says "I have a delivery," and the voice is captured by the device and sent to the server.
[0459] Step 7:
[0460] The server analyzes what the visitor says and generates an appropriate response.
[0461] Input: Audio data sent from the device
[0462] Data processing: Converting voice data into text using Google Cloud Speech-to-Text API, etc., and then generating an appropriate response.
[0463] Output: Appropriate response text (e.g., "Okay, you can pick it up with the QR code.")
[0464] What happens: The server converts the voice data into text and generates an appropriate response based on that text.
[0465] Step 8:
[0466] The server adjusts the digital human's response based on the emotion analysis results.
[0467] Input: Response text and sentiment data
[0468] Data processing: Adjust tone and facial expression based on emotion analysis results
[0469] Output: Adjusted response data
[0470] Specific behavior: The server responds with a gentler tone, "Okay, you can pick it up with the QR code."
[0471] Step 9:
[0472] The server generates a QR code, which the device displays.
[0473] Input: Visitor's payment and receipt instructions
[0474] Data processing: Generate a QR code using the Python qrcode library
[0475] Output: Generated QR code data
[0476] Specific operation: When a visitor requests receipt using a QR code, the server generates a QR code and sends it to the terminal, which displays it.
[0477] Step 10:
[0478] The server records all conversations and videos, and the user can check the recorded content.
[0479] Input: Dialogue content, video, emotion data
[0480] Data processing: Recording and saving in a database
[0481] Output: Recorded dialogue and video data
[0482] Specific operation: All conversations and videos are recorded on the server and can be viewed by the user through a dedicated interface.
[0483] (Application example 2)
[0484] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0485] Conventional customer service systems have difficulty recognizing customers' emotions and responding appropriately, limiting the ways to improve customer satisfaction. Furthermore, they lack the ability to respond based on the visitor's attributes and emotions, meaning that even customers who require special attention can only receive a uniform level of service. As a result, customers often become dissatisfied.
[0486] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0487] In this invention, the server includes means for analyzing video data to identify the visitor's attributes and emotions, means for generating video and audio of a digital human based on the visitor's attributes and emotions, and means for interacting with the visitor and adjusting responses based on the emotions, thereby enabling flexible responses according to the customer's emotions.
[0488] "Camera Device" means a device used to capture video of a visitor.
[0489] "Video data" refers to video information captured by a camera device.
[0490] The "server" is a device that analyzes video data, identifies the visitor's attributes and emotions, and generates video and audio of a digital human.
[0491] "Attributes" refer to personal characteristics of visitors, such as their age, gender, and belongings.
[0492] "Emotion" refers to the visitor's mental state, and includes states such as joy, anger, and anxiety.
[0493] A "digital human" is a virtual human image and voice generated to interact with visitors.
[0494] "Dialogue" refers to audio and video communication between the digital human and the visitor.
[0495] "Recording" means saving the content of conversations and footage of visitors.
[0496] A QR code is a type of barcode that visitors can scan with a smartphone or other device to receive information or make payments.
[0497] "Audio Recording Device" means a device used to capture the audio of a visitor.
[0498] "Analysis" is the act of processing received data and extracting information.
[0499] A system for implementing the present invention mainly includes the following components:
[0500] 1. Camera equipment:
[0501] The camera device is installed to capture video of visitors in real time. Any commercially available video camera can be used as the camera device. For example, the Logitech C920 can be used.
[0502] 2. Server:
[0503] The server receives the video data sent from the camera and performs the necessary analysis to identify the visitor's attributes and emotions. Specifically, it uses "OpenCV + Dlib" to analyze the video data, and "Microsoft Azure Cognitive Services" can be used for emotion recognition. The server also uses a combination of "Unity + speech synthesis technology" to generate the video and audio of the digital human.
[0504] 3. Terminal:
[0505] The terminal is a device that receives data sent from the server and enables interaction with visitors. The terminal must be equipped with a voice recording device to capture the visitor's voice and speech recognition technology. For this purpose, "Google Speech-to-Text" can be used. Also, the "Zebra Crossing (ZXing)" library can be used to generate QR codes.
[0506] 4. Emotion Engine:
[0507] The emotion engine is used to identify the emotional state of the visitor by analyzing their video and audio data, and can use emotion recognition technologies such as those from Microsoft Azure Cognitive Services.
[0508] 5. User Interface:
[0509] The user interface is an interface that allows users to configure the system and check records. A web-based user interface can be built using "React.js."
[0510] To give a specific example, the following system is realized.
[0511] Example 1
[0512] When a visitor enters a store, a camera captures the visitor's video and sends it to a server. The server analyzes the video and identifies the visitor's attributes (age, gender, etc.) and emotions. For example, if the server determines that the visitor is a man in his 40s who is feeling somewhat stressed, it generates a digital human based on that information and displays it on the terminal. The terminal then begins a dialogue with the visitor through the digital human, responding flexibly according to the visitor's emotions. If the visitor wishes to purchase an item and selects QR code payment, the terminal displays the QR code received from the server, and the payment is completed by having the visitor scan it with their smartphone.
[0513] Examples of prompt sentences include:
[0514] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[0515] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0516] Step 1:
[0517] Capture footage of visitors.
[0518] The camera device captures video of visitors in real time and transmits the video data to the server. The input data is raw video data, and the output is video data transmitted to the server. Specifically, the camera device continuously captures video and transmits the data to the server via the network.
[0519] Step 2:
[0520] Analyzing video data and identifying visitor attributes and emotions.
[0521] The server analyzes the transmitted video data and identifies the visitor's attributes (age, gender, etc.). It then uses an emotion engine to analyze the visitor's emotions. The input data is the video data, and the output data is the visitor's attributes and emotional information. Specifically, the server first performs face detection and feature extraction using "OpenCV + Dlib," and then performs emotion analysis on the results using "Microsoft Azure Cognitive Services."
[0522] Step 3:
[0523] Digital human generation.
[0524] The server generates video and audio of a digital human based on the identified attributes and emotions. The input data is the visitor's attributes and emotional information, and the output data is the video and audio data of the generated digital human. Specifically, the video of the digital human is generated using "Unity," and the audio is generated using "voice synthesis technology."
[0525] Step 4:
[0526] Digital human interacting with visitors.
[0527] The terminal displays the digital human received from the server and begins a dialogue with the visitor. The input data is the video and audio data of the digital human, and the output data is the content of the dialogue with the visitor. Specifically, the digital human is displayed on the terminal's display and audio is played back through the audio speaker.
[0528] Step 5:
[0529] Recording and analyzing visitor voices.
[0530] The device records the visitor's voice, analyzes the voice data, and generates an appropriate response. The input data is the visitor's voice data, and the output data is the analyzed voice text and the generated response. Specifically, the device's built-in microphone captures the voice, converts it to text using Google Speech-to-Text, and then the server analyzes the text to generate an appropriate response.
[0531] Step 6:
[0532] Generate and display QR codes.
[0533] When a visitor requests QR code payment, the server generates a dedicated QR code and sends it to the terminal. The terminal then displays it to the visitor. The input data is the visitor's payment preference information, and the output data is the generated QR code. Specifically, the server uses the "Zebra Crossing (ZXing)" library to generate the QR code and display it on the terminal.
[0534] Step 7:
[0535] Recording of conversations and video footage.
[0536] The server records and saves the conversations and video data with visitors. The input data is the conversations and video data, and the output data is the saved information. Specifically, the server uses cloud storage such as AWS S3 to save the data so that it can be viewed later.
[0537] An example of a prompt sentence would be something like:
[0538] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[0539] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0540] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0541] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0542] [Second embodiment]
[0543] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0544] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0545] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0546] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0547] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0548] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0549] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0550] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0551] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0552] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0553] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0554] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0555] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[0556] System Configuration
[0557] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[0558] 2. Server: Receives and analyzes the video data sent from the camera device, identifies visitor attributes, and generates video and audio of a digital human based on those attributes.
[0559] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[0560] 4. User Interface: This is the interface that allows users to configure the system and check records.
[0561] Program processing explanation
[0562] 1. Visitor Recognition
[0563] The terminal transmits the video data from the camera device to the server.
[0564] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0565] 2. Digital Human Generation
[0566] The server generates the appropriate video and audio of the digital human based on the identified attributes.
[0567] For example, the setting is such that elderly people speak in a slow tone, while children speak in a friendly tone.
[0568] 3. Interacting with visitors
[0569] The terminal displays the digital human received from the server and greets the visitor.
[0570] The visitor speaks to the digital human.
[0571] The terminal picks up the visitor's voice and transmits it to the server.
[0572] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[0573] The terminal communicates the generated answer to the visitor via a digital human.
[0574] 4. QR Code Processing
[0575] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[0576] The server generates a QR code and sends it to the device.
[0577] The terminal displays a QR code that visitors can scan with their smartphone.
[0578] 5. Recording and Confirmation
[0579] The server records all conversations and videos.
[0580] Users can later review the recorded conversations and footage through a dedicated interface.
[0581] Specific examples
[0582] Visitor recognition and digital human generation
[0583] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[0584] Interacting with visitors
[0585] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates the appropriate response, "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[0586] QR code processing and recording
[0587] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[0588] The above is an embodiment of the present invention. The present invention automates the process of dealing with visitors, significantly reducing the effort and risk for users.
[0589] The processing flow will be explained below.
[0590] Step 1:
[0591] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[0592] Step 2:
[0593] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[0594] Step 3:
[0595] The server generates the appropriate video and audio of a digital human based on the identified attributes, for example, an elderly person will have a slower voice.
[0596] Step 4:
[0597] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[0598] Step 5:
[0599] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[0600] Step 6:
[0601] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[0602] Step 7:
[0603] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[0604] Step 8:
[0605] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[0606] Step 9:
[0607] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[0608] Step 10:
[0609] The visitor tells the digital human that they would like to process the transaction using a QR code.
[0610] Step 11:
[0611] The server generates a QR code and sends it to the device.
[0612] Step 12:
[0613] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[0614] Step 13:
[0615] The server records all interactions and videos with visitors and stores them in a database.
[0616] Step 14:
[0617] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[0618] Example 1
[0619] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0620] There is a need to build a system that automates visitor interactions and provides appropriate responses based on the visitor's attributes. Conventional visitor interaction systems have difficulty accurately identifying attributes such as the visitor's age and gender and generating a suitable digital human. Furthermore, many systems lack the ability to record and review the content and video of interactions with visitors. Furthermore, there is a lack of a convenient way to make payments and receive items using QR codes.
[0621] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0622] In this invention, the server includes means for analyzing video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, and means for generating a greeting for the digital human using a generative AI model. This makes it possible to generate and display an appropriate digital human according to the visitor attributes, automatically generate a greeting and response according to the attributes, and record and check the content of conversations with visitors and their video.
[0623] A "camera device" is a photographic device for capturing images of visitors, and is installed outside the door, etc.
[0624] "Video data" refers to video information of visitors captured by a camera device.
[0625] "Analyzing video data" means performing processing to identify visitor attribute information (age, gender, belongings, etc.) from the received video data.
[0626] "Visitor attributes" refers to characteristic information about a visitor, such as age, gender, and belongings.
[0627] "Digital Human" means a virtual character, including video and audio, that is generated to interact with visitors.
[0628] An "audio recording device" is a device such as a microphone for capturing the visitor's voice.
[0629] "Voice data" refers to audio information that records the speech of a visitor.
[0630] "Speech recognition technology" is a technology for converting voice data into text and analyzing the content of speech.
[0631] A "generative AI model" is an artificial intelligence model that generates appropriate responses and greetings based on attribute information and speech content.
[0632] A "QR code" is a two-dimensional barcode that visitors use when making payments or receiving items.
[0633] "Recording" refers to the act of saving the content of conversations with visitors and footage.
[0634] "User interface" refers to the interface that allows the user to configure the system and check the recorded dialogue and video.
[0635] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[0636] System Configuration
[0637] 1. Camera device: Installed on the exterior of the door to capture video of visitors. It has the ability to capture high-resolution video, capturing detailed images of visitors' faces and belongings.
[0638] 2. Server: Receives and analyzes the video data sent from the camera. The server incorporates image analysis software (e.g., OpenCV or TensorFlow) and speech recognition technology (e.g., Google Speech-to-Text API). It also generates digital human responses using generative AI models (e.g., OpenAI's GPT-3.5).
[0639] 3. Terminal: Receives data from the server and interacts with visitors. The terminal is equipped with a voice recording device that captures the visitor's voice and sends it to the server.
[0640] 4. User Interface: This is the interface through which users can configure the system, view recorded conversations, and check video footage. It is built as a web application and can be accessed from a browser (e.g., using React or Angular).
[0641] Program processing explanation
[0642] 1. Visitor Recognition:
[0643] The device sends the video data from the camera to the server, which analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0644] 2. Digital Human Generation:
[0645] The server then generates the appropriate video and audio of a digital human based on the identified attributes, and the generated digital human is configured to speak in different tones depending on the visitor's attributes, for example, speaking in a slower tone for elderly people and in a more friendly tone for children.
[0646] 3. Visitor interaction:
[0647] The device displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the device's audio pickup device and sent to the server. The server uses speech recognition technology to analyze the speech and generates an appropriate response using a generative AI model. The device then conveys this response to the visitor via the digital human.
[0648] 4. QR Code Processing:
[0649] When a visitor wishes to make or receive payment using a QR code, they tell the digital human, and the server generates a QR code and sends it to the terminal, which displays the code and the visitor scans it with their smartphone.
[0650] 5. Record and verify:
[0651] The server records all conversations and videos, and users can later review the recorded conversations and videos through a dedicated interface.
[0652] Specific examples
[0653] Visitor recognition and digital human generation
[0654] One day, an elderly visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model designed for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[0655] Interacting with visitors
[0656] The visitor says, "You have a delivery," and the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and generates an appropriate response based on the text: "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[0657] QR code processing and recording
[0658] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[0659] Prompt Sentence Examples
[0660] Here are some example prompts to input to a generative AI model (e.g., OpenAI's GPT-3.5):
[0661] Generate digital human greetings based on visitor attributes. Digital human greetings based on the following attributes:
[0662] Age: Elderly
[0663] Gender: Female
[0664] Possession: Cane
[0665] Please speak in a relaxed, friendly tone for seniors.
[0666] Using this prompt, the system is capable of generating an appropriate digital human greeting for the visitor.
[0667] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0668] Step 1:
[0669] Video capture using camera equipment
[0670] The terminal acquires video data from a camera device installed outside the door.
[0671] Input: Real-time video of the visitor captured by the camera device.
[0672] Output: Captured video data.
[0673] Step 2:
[0674] Video data sent to server
[0675] The terminal compresses the acquired video data, minimizes the amount of data, and transmits it to the server.
[0676] Input: Video data acquired from a camera device.
[0677] Output: Compressed video data sent to the server.
[0678] Step 3:
[0679] Identifying visitor attributes
[0680] The server uses image analysis software (e.g., OpenCV or TensorFlow) to analyze the received video data.
[0681] Through video analysis, the server identifies visitors' attributes such as age, gender, and belongings.
[0682] Input: Compressed video data.
[0683] Output: Visitor demographic information (e.g. age, gender, belongings).
[0684] Step 4:
[0685] Digital Human Generation
[0686] The server uses a generative AI model (e.g., OpenAI's GPT-3.5) to generate a greeting for the digital human based on the visitor's identified attribute information.
[0687] Next, 3D modeling software (e.g., Blender or Unity) is used to generate the video and audio of the digital human.
[0688] Input: Visitor demographic information.
[0689] Output: Digital human video and audio data.
[0690] Step 5:
[0691] Digital human display and greeting
[0692] The terminal displays the video and audio of the digital human received from the server.
[0693] A digital human greets visitors.
[0694] Input: Video and audio data of a digital human.
[0695] Output: A displayed digital human greeting.
[0696] Step 6:
[0697] Interacting with visitors
[0698] The visitor speaks to the digital human.
[0699] The terminal picks up the visitor's speech and sends the audio data to the server.
[0700] Input: Visitor utterance.
[0701] Output: The audio data sent to the server.
[0702] Step 7:
[0703] Voice data analysis and response generation
[0704] The server uses voice recognition technology (e.g., Google Speech-to-Text API) to convert the visitor's speech into text.
[0705] It then uses a generative AI model to generate appropriate responses to what the visitor says.
[0706] Input: Visitor's voice data.
[0707] Output: The recognized text and the response message.
[0708] Step 8:
[0709] Viewing the response
[0710] The terminal conveys the response message received from the server to the visitor through the digital human.
[0711] Input: Digital human's response message.
[0712] Output: Displayed digital human response.
[0713] Step 9:
[0714] QR code processing
[0715] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[0716] The server generates a dedicated QR code and sends it to the device.
[0717] The terminal displays a QR code that visitors can scan with their smartphone.
[0718] Input: Visitor processing request.
[0719] Output: The displayed QR code.
[0720] Step 10:
[0721] Dialogue and video recording
[0722] The server records all conversations and videos.
[0723] Input: Dialogue content and video data.
[0724] Output: Recorded dialogue and video data.
[0725] Step 11:
[0726] Checking the recorded content
[0727] Users can view the recorded conversations and footage through a dedicated user interface.
[0728] Input: Request to access recorded data.
[0729] Output: Displayed recording and video data.
[0730] (Application example 1)
[0731] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0732] Conventional visitor response systems have had several limitations in terms of visitor recognition, interaction, and enhanced security. Specifically, they have difficulty in flexibly interacting with visitors and responding quickly, which increases the workload of users. Furthermore, they have been unable to provide appropriate responses based on visitor attributes or automatically manage records. This has led to problems with systems not functioning effectively, particularly when individualized attention is required for elderly people and children.
[0733] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0734] In this invention, the server includes means for capturing video of nearby visitors with a camera device, means for analyzing the video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, means for interacting with the visitor, means for recording the content of the interaction and the visitor video, means for interacting with the visitor using a smartphone, means for capturing the visitor's audio with the smartphone and analyzing the audio data to generate an appropriate response, and means for generating and displaying a QR code on the smartphone. This enables flexible visitor recognition and appropriate response, reduces user effort, and enhances security.
[0735] A "camera device" is a device for capturing video of visitors in real time.
[0736] "Video data" refers to video information of visitors captured by a camera device.
[0737] "Visitor attributes" refers to information such as the visitor's age, gender, belongings, etc.
[0738] A "digital human" is a video and audio representation of a virtual person generated based on the visitor's attributes.
[0739] "Dialogue" refers to communication between a digital human and a visitor.
[0740] "Recording" refers to saving the content of conversations and footage of visitors.
[0741] A "smartphone" is a portable information terminal that works in conjunction with a camera device and a server and is used to interact with visitors.
[0742] "Audio data" refers to information that records the content of a visitor's speech.
[0743] A "QR code" is a two-dimensional barcode generated for visitors to make payments and other transactions.
[0744] The "server" is a central processing unit that receives data from the camera device, analyzes it, generates digital humans, and records them.
[0745] MODE FOR CARRYING OUT THE INVENTION
[0746] A system for implementing the present invention is configured as follows: First, a camera device captures video of nearby visitors in real time and transmits the video data to a server. Next, the server analyzes the video data and identifies the visitor's attributes. Based on the identified attributes, the server generates video and audio of a digital human and interacts with the visitor via a smartphone.
[0747] Required Hardware and Software
[0748] Camera equipment: Equipment for capturing images of visitors. High-resolution cameras are recommended.
[0749] Server: A central processing unit that analyzes data, generates digital humans, and records data. It incorporates machine learning libraries such as TensorFlow and PyTorch.
[0750] Smartphone: A mobile information device for interacting with visitors. It requires an internet connection, a camera, and a microphone.
[0751] Speech recognition and speech synthesis technologies: The server incorporates speech recognition technologies (such as the Google Speech-to-Text API) and speech synthesis technologies (such as the Google Text-to-Speech API).
[0752] QR code generation software: A library for generating QR codes (e.g., Python's QRCode library).
[0753] Processing flow
[0754] Step 1: Video capture
[0755] The camera captures video of visitors in real time and sends the video data to a server, which analyzes the received video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0756] Step 2: Creating a digital human
[0757] The server generates the appropriate video and audio of the digital human based on the identified attributes. For example, it may be configured to speak in a slower tone for an elderly person, or in a more friendly tone for a child. The data of the generated digital human is then sent to a smartphone.
[0758] Step 3: Dialogue with visitors
[0759] The smartphone displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the smartphone's microphone and sent to the server. The server uses voice recognition technology to analyze the speech and generate an appropriate response. The smartphone then broadcasts the generated response via the digital human.
[0760] Step 4: Processing by QR code
[0761] When a visitor wishes to make a payment or receive payment using a QR code, they notify the digital human. The server generates a dedicated QR code and sends it to the smartphone. The smartphone displays this QR code, and the visitor can scan it to complete the payment process.
[0762] Step 5: Record and verify
[0763] The server records all conversations and video footage, which users can later review through a dedicated interface. This automates interactions with visitors, significantly reducing the user's workload and risk.
[0764] Examples and prompts
[0765] Specific examples
[0766] One day, a visitor approaches the door of a home. A camera installed on the door captures the visitor's video and sends it to a server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people, generates video and audio of the visitor speaking in a slow tone, saying, "Hello. How can I help you?" and sends this video and audio to the smartphone. The smartphone then displays this digital human and greets the visitor.
[0767] Prompt Sentence Examples
[0768] "Generate the appropriate digital human based on the visitor's attributes. For example, set a slower tone for seniors and a more friendly tone for children. Also, generate a QR code."
[0769] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0770] Step 1:
[0771] Input: Video data captured by a camera device
[0772] Processing: The camera device captures video of nearby visitors in real time and collects the video data.
[0773] Output: Collected video data
[0774] Specific operation: The camera captures the moment a visitor approaches the door and saves the video data.
[0775] Step 2:
[0776] Input: Video data obtained from a camera device
[0777] Processing: The server receives the video data and uses video analytics techniques (e.g., facial and object recognition) to identify visitor attributes (e.g., age, gender, belongings, etc.).
[0778] Output: Visitor attribute data identified by analysis
[0779] Specific operation: The server analyzes the video data using TensorFlow and PyTorch, identifies the visitor's face and features, and determines their age and gender.
[0780] Step 3:
[0781] Input: Visitor attribute data
[0782] Processing: The server generates the appropriate video and audio of the digital human based on the visitor's attributes.
[0783] Output: Video and audio data of the generated digital human
[0784] How it works: The server uses the generative AI model to generate a digital human that, for example, greets elderly people in a slower tone. Video and audio data are generated.
[0785] Step 4:
[0786] Input: Digital human video and audio data
[0787] Processing: The server sends the generated digital human data to the smartphone, which displays the digital human to the visitor and greets them.
[0788] Output: Video and audio of a digital human displayed on a smartphone
[0789] Specific operation: The smartphone plays the video of the digital human it receives and greets the user with a greeting such as "Hello. How can I help you?"
[0790] Step 5:
[0791] Input: What the visitor said
[0792] Processing: The device picks up what the visitor is saying and sends the audio data to the server, which uses speech recognition technology to convert it into text data and generate an appropriate response.
[0793] Output: Speech-to-text data and the digital human's response based on it
[0794] Specific operation: The smartphone uses a microphone to pick up the visitor's speech, such as "I have a delivery for you," and the server converts this into text data and generates a response such as "Okay, you can pick it up using the QR code."
[0795] Step 6:
[0796] Input: Visitor's request (e.g., request for pickup via QR code)
[0797] Process: The server generates a QR code and sends it to the smartphone, which displays the generated QR code to the visitor.
[0798] Output: QR code displayed on the smartphone
[0799] How it works: The server generates a QR code using a Python QRCode library or similar, and the smartphone displays the QR code to the visitor.
[0800] Step 7:
[0801] Input: Dialogue content and video data
[0802] Processing: The server records all conversations and video. The user can view the recorded data through a dedicated interface.
[0803] Output: Recorded dialogue and video data
[0804] How it works: The server stores all the interaction data and footage in a database, allowing users to view this data later.
[0805] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0806] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, an emotion engine, and a user interface.
[0807] System Configuration
[0808] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[0809] 2. Server: Receives and analyzes the video and audio data sent from the camera device, identifies the visitor's attributes and emotions, and generates the video and audio of a digital human based on those.
[0810] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[0811] 4. Emotion Engine: An engine that analyzes the visitor's video and audio data and recognizes the visitor's emotions.
[0812] 5. User Interface: This is the interface that allows users to configure the system and check records.
[0813] Program processing explanation
[0814] 1. Visitor Recognition
[0815] The terminal transmits the video data from the camera device to the server.
[0816] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[0817] 2. Emotional Recognition
[0818] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[0819] Identify different emotional states of your visitors, such as happy, angry, or anxious.
[0820] 3. Digital Human Generation
[0821] The server generates the appropriate video and audio of the digital human based on the identified attributes and emotions.
[0822] For example, if an elderly visitor is feeling anxious, you might adopt a more friendly and gentle tone.
[0823] 4. Interacting with visitors
[0824] The terminal displays the digital human received from the server and makes initial responses to visitors, such as greetings.
[0825] Visitors speak to the digital human.
[0826] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[0827] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[0828] 5. Regulating responses based on emotions
[0829] Based on the analysis results of the emotion engine, the server adjusts the visual and audio tone of the digital human, for example, responding in a calmer tone to soothe the visitor's anger.
[0830] 6. QR Code Processing
[0831] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[0832] The server generates a QR code and sends it to the device.
[0833] The terminal displays a QR code that visitors can scan with their smartphone.
[0834] 7. Recording and Confirmation
[0835] The server records all conversations and videos.
[0836] The recording also includes visitor sentiment data, which users can review later.
[0837] Users can later review the recorded conversations and footage through a dedicated interface.
[0838] Specific examples
[0839] Visitor recognition and emotion recognition
[0840] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. It also uses an emotion engine to identify that the visitor is feeling anxious. The server then generates a digital human with a gentle tone of voice, specifically for the elderly, to ease their anxiety.
[0841] Interact with visitors and tailor responses based on their emotions
[0842] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates an appropriate response, "Okay, you can pick it up by scanning the QR code." Based on the analysis results of the emotion engine, the response is adjusted to a gentler tone. The device then plays this response on a digital human and conveys it to the visitor.
[0843] QR code processing and recording
[0844] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the device. The device displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server and saved, including emotional data. Users can later review these records through a dedicated interface.
[0845] The above is an embodiment of the present invention. This invention automates the process of interacting with visitors, significantly reducing the user's workload and risk. By combining it with an emotion engine, it becomes possible to respond appropriately to visitors' emotional states.
[0846] The processing flow will be explained below.
[0847] Step 1:
[0848] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[0849] Step 2:
[0850] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[0851] Step 3:
[0852] The server analyzes the visitor's emotions from the video and audio data using an emotion engine, which identifies the visitor's emotional state from their facial expressions and tone of voice.
[0853] Step 4:
[0854] The server generates an appropriate digital human image and voice based on the identified attributes and emotions. For example, a digital human with a gentle tone is generated for an elderly visitor who is anxious.
[0855] Step 5:
[0856] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[0857] Step 6:
[0858] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[0859] Step 7:
[0860] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[0861] Step 8:
[0862] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[0863] Step 9:
[0864] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[0865] Step 10:
[0866] Based on the analysis results of the emotion engine, the server adjusts the video and audio tones of the digital human and responds according to the visitor's emotional state.
[0867] Step 11:
[0868] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[0869] Step 12:
[0870] The visitor tells the digital human that they would like to process the transaction using a QR code.
[0871] Step 13:
[0872] The server generates a QR code and sends it to the device.
[0873] Step 14:
[0874] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[0875] Step 15:
[0876] The server records all interactions and videos with visitors and stores them in a database, including emotional data analyzed by the emotion engine.
[0877] Step 16:
[0878] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[0879] Example 2
[0880] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0881] Automated response systems that interact with visitors are required to accurately recognize the visitor's attributes and emotions and generate an appropriate digital human. It is also important to enable visitors to smoothly make payments and receive payments using QR codes. However, integrating these elements into a single system and making it function smoothly is technically difficult.
[0882] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0883] In this invention, the server includes means for capturing video of nearby visitors using a camera device, means for analyzing the video data and identifying the visitor's attributes, means for analyzing the visitor's emotions from the video and audio data, means for generating video and audio of a digital human based on the visitor's attributes and emotions, means for displaying the digital human and engaging in a dialogue with the visitor, means for capturing the visitor's audio and generating an appropriate response using voice recognition technology, means for adjusting the digital human's response based on the emotion analysis results, means for generating and displaying a QR code when the visitor uses a QR code to make or receive payments, and means for recording the content of the dialogue and the visitor's video. This allows the visitor's attributes and emotions to be accurately recognized, enabling smooth payment and receipt using an appropriate digital human for dialogue and the QR code.
[0884] "Camera Device" means equipment used to capture real-time video of visitors.
[0885] "Video Data" refers to still and video images of Visitors captured by camera devices.
[0886] The "server" is a central processing unit that analyzes video and audio data and identifies the attributes and emotions of visitors.
[0887] "Visitor attributes" refers to characteristic information about visitors, such as age, gender, and belongings.
[0888] The "Emotion Engine" is a software interface for analyzing a visitor's emotional state from video and audio data.
[0889] A "digital human" is an artificial visual and audio character that is generated to interact with visitors.
[0890] "Voice recognition technology" is a technology for converting a visitor's voice into text data.
[0891] A "QR code" is a two-dimensional barcode used by visitors to make payments and receive goods.
[0892] "Dialogue content" refers to the entire content of the conversation between the visitor and the digital human.
[0893] "Recording means" refers to methods and devices for saving the content of interactions and video of visitors.
[0894] In one embodiment of the present invention, the system is composed of a camera device, a server, a terminal, an emotion engine, and a user interface, which allows the system to recognize the attributes and emotions of visitors and generate an appropriate digital human to interact with them.
[0895] System Configuration
[0896] camera equipment
[0897] The camera device is installed outside the door and captures visitors' images in real time. The camera device acquires high-resolution images and transmits them to a terminal.
[0898] server
[0899] The server receives and analyzes the video and audio data sent from the camera. The server has the following functions:
[0900] 1. Video analysis: Using a video analysis library such as OpenCV, we identify visitors' faces and extract attributes such as age, gender, and belongings.
[0901] 2. Emotion analysis: Using Microsoft Azure's emotion recognition API, we analyze visitors' emotions from video and audio data.
[0902] 3. Digital Human Generation: Using tools such as Adobe Character Animator, generate the video and audio of a digital human based on identified attributes and emotions.
[0903] 4. Speech Recognition: Using APIs such as Google Cloud Speech-to-Text, convert the visitor's voice data into text and generate an appropriate response.
[0904] 5. Response adjustment: Adjust the digital human's response based on the emotion analysis results.
[0905] Terminal
[0906] The terminal is the device that receives data from the server and interacts with the visitor. The terminal has the following functions:
[0907] 1. Video display: Displays the digital human received from the server and provides initial greetings and guidance to visitors.
[0908] 2. Microphone: A built-in microphone is included to capture the visitor's speech and send it to the server.
[0909] 3. Display QR code: When a visitor uses a QR code to make or receive a payment, the QR code sent from the server is displayed.
[0910] Emotion Engine
[0911] The emotion engine analyzes the video and audio data of visitors and recognizes their emotions. This engine can identify their emotional state, such as whether they are happy, angry, or anxious.
[0912] User Interface
[0913] Users can use the interface to configure the system and view recordings, which include visitor attributes, emotional state, dialogue, and video.
[0914] Specific operation example
[0915] 1. Recognizing visitors' perceptions and emotions
[0916] One day, a visitor approaches the door of your home. A camera installed on the door captures the visitor's image, and the device sends it to the server.
[0917] The server uses video analysis to determine that the visitor is elderly, and an emotion engine to identify that the visitor is feeling anxious.
[0918] The server is designed for the elderly, generating a digital human with a gentle tone to ease anxiety.
[0919] 2. Interact with visitors and tailor responses based on their emotions
[0920] When a visitor says "I have a delivery," the voice is captured by the device's microphone and sent to the server.
[0921] The server uses speech recognition technology to generate the text "You have a delivery" and then generates the appropriate response "Okay, you can pick it up using the QR code."
[0922] Based on the analysis results of the emotion engine, the response is adjusted to a gentle tone, which the device then plays back to the digital human and conveys to the visitor.
[0923] 3. Processing and recording by QR code
[0924] When a visitor requests to receive their gift using a QR code, the server generates a dedicated QR code and sends it to the device.
[0925] The terminal displays this QR code, and visitors can scan it with their smartphone to complete the pickup.
[0926] All conversations and videos are recorded on a server, including emotional data, and users can view these records through a dedicated interface.
[0927] Prompt Sentence Examples
[0928] 1. Prompt to identify visitor demographics:
[0929] Identify the visitor's age, gender, and belongings from the video data, for example, whether the visitor is carrying a bag.
[0930] 2. Emotion analysis prompts:
[0931] Analyze visitor emotions from this video and audio data. Identify whether your visitors are happy, angry, anxious, etc.
[0932] 3. Prompt when generating a digital human:
[0933] If your visitor is elderly and anxious, generate a digital human that responds in a friendly and gentle tone.
[0934] 4. Speech recognition prompts:
[0935] Convert the visitor's speech into text data. For example, convert the speech "I have a delivery" into text.
[0936] The above is an embodiment of the present invention. The present invention enables accurate recognition of visitor attributes and emotions, and enables smooth payment and receipt using appropriate digital human dialogue and QR codes.
[0937] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0938] Step 1:
[0939] The terminal acquires video data from the camera device and transmits it to the server.
[0940] Input: Real-time video data from a camera device
[0941] Data processing: The device reads the video data from the buffer.
[0942] Output: Video data transmitted to the server through the configured network connection
[0943] Specific operation: A camera device is installed outside the door, and when a visitor approaches, it automatically captures video and the device sends the video data to the server.
[0944] Step 2:
[0945] The server analyzes the received video data and identifies the visitor's attributes.
[0946] Input: Video data sent from the device
[0947] Data processing: Using a video analysis library such as OpenCV, a facial recognition algorithm is applied to extract attributes such as age, gender, and belongings.
[0948] Output: Visitor attribute data (e.g., age: 60, gender: male, belongings: bag)
[0949] What happens: The server performs video analysis to detect the visitor's face and identify their attributes.
[0950] Step 3:
[0951] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[0952] Input: Analyzed video and audio data
[0953] Data processing: Analyze facial expressions and tone of voice using Microsoft Azure's emotion recognition API.
[0954] Output: Visitor emotion data (e.g., anxiety, joy, anger)
[0955] What happens: The server sends data to the emotion engine to identify the visitor's emotion.
[0956] Step 4:
[0957] The server generates a digital human based on the visitor's identified attributes and emotions.
[0958] Input: Visitor demographic and sentiment data
[0959] Data processing: Using Adobe Character Animator or similar software, we generate appropriate digital human images and voices.
[0960] Output: Video and audio data of the generated digital human
[0961] Specific behavior: The server generates a digital human with a friendly and gentle tone that corresponds to the identified attributes (e.g., elderly, anxious state).
[0962] Step 5:
[0963] The terminal displays the digital human received from the server and makes an initial response to the visitor.
[0964] Input: Video and audio data of the digital human sent from the server
[0965] Data processing: Display and play data received by the terminal
[0966] Output: The video and audio of the digital human are displayed and played back to the visitor.
[0967] What it does: The device greets the visitor with the video and audio of a digital human, asking, "Hello, how can I help you?"
[0968] Step 6:
[0969] The visitor speaks to the digital human, and the device captures the audio and sends it to the server.
[0970] Input: Visitor's speech
[0971] Data processing: The device's microphone captures the audio and sends the audio data to the server.
[0972] Output: Audio data sent to the server
[0973] Specific operation: The visitor says "I have a delivery," and the voice is captured by the device and sent to the server.
[0974] Step 7:
[0975] The server analyzes what the visitor says and generates an appropriate response.
[0976] Input: Audio data sent from the device
[0977] Data processing: Converting voice data into text using Google Cloud Speech-to-Text API, etc., and then generating an appropriate response.
[0978] Output: Appropriate response text (e.g., "Okay, you can pick it up with the QR code.")
[0979] What happens: The server converts the voice data into text and generates an appropriate response based on that text.
[0980] Step 8:
[0981] The server adjusts the digital human's response based on the emotion analysis results.
[0982] Input: Response text and sentiment data
[0983] Data processing: Adjust tone and facial expression based on emotion analysis results
[0984] Output: Adjusted response data
[0985] Specific behavior: The server responds with a gentler tone, "Okay, you can pick it up with the QR code."
[0986] Step 9:
[0987] The server generates a QR code, which the device displays.
[0988] Input: Visitor's payment and receipt instructions
[0989] Data processing: Generate a QR code using the Python qrcode library
[0990] Output: Generated QR code data
[0991] Specific operation: When a visitor requests receipt using a QR code, the server generates a QR code and sends it to the terminal, which displays it.
[0992] Step 10:
[0993] The server records all conversations and videos, and the user can check the recorded content.
[0994] Input: Dialogue content, video, emotion data
[0995] Data processing: Recording and saving in a database
[0996] Output: Recorded dialogue and video data
[0997] Specific operation: All conversations and videos are recorded on the server and can be viewed by the user through a dedicated interface.
[0998] (Application example 2)
[0999] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1000] Conventional customer service systems have difficulty recognizing customers' emotions and responding appropriately, limiting the ways to improve customer satisfaction. Furthermore, they lack the ability to respond based on the visitor's attributes and emotions, meaning that even customers who require special attention can only receive a uniform level of service. As a result, customers often become dissatisfied.
[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1002] In this invention, the server includes means for analyzing video data to identify the visitor's attributes and emotions, means for generating video and audio of a digital human based on the visitor's attributes and emotions, and means for interacting with the visitor and adjusting responses based on the emotions, thereby enabling flexible responses according to the customer's emotions.
[1003] "Camera Device" means a device used to capture video of a visitor.
[1004] "Video data" refers to video information captured by a camera device.
[1005] The "server" is a device that analyzes video data, identifies the visitor's attributes and emotions, and generates video and audio of a digital human.
[1006] "Attributes" refer to personal characteristics of visitors, such as their age, gender, and belongings.
[1007] "Emotion" refers to the visitor's mental state, and includes states such as joy, anger, and anxiety.
[1008] A "digital human" is a virtual human image and voice generated to interact with visitors.
[1009] "Dialogue" refers to audio and video communication between the digital human and the visitor.
[1010] "Recording" means saving the content of conversations and footage of visitors.
[1011] A QR code is a type of barcode that visitors can scan with a smartphone or other device to receive information or make payments.
[1012] "Audio Recording Device" means a device used to capture the audio of a visitor.
[1013] "Analysis" is the act of processing received data and extracting information.
[1014] A system for implementing the present invention mainly includes the following components:
[1015] 1. Camera equipment:
[1016] The camera device is installed to capture video of visitors in real time. Any commercially available video camera can be used as the camera device. For example, the Logitech C920 can be used.
[1017] 2. Server:
[1018] The server receives the video data sent from the camera and performs the necessary analysis to identify the visitor's attributes and emotions. Specifically, it uses "OpenCV + Dlib" to analyze the video data, and "Microsoft Azure Cognitive Services" can be used for emotion recognition. The server also uses a combination of "Unity + speech synthesis technology" to generate the video and audio of the digital human.
[1019] 3. Terminal:
[1020] The terminal is a device that receives data sent from the server and enables interaction with visitors. The terminal must be equipped with a voice recording device to capture the visitor's voice and speech recognition technology. For this purpose, "Google Speech-to-Text" can be used. Also, the "Zebra Crossing (ZXing)" library can be used to generate QR codes.
[1021] 4. Emotion Engine:
[1022] The emotion engine is used to identify the emotional state of the visitor by analyzing their video and audio data, and can use emotion recognition technologies such as those from Microsoft Azure Cognitive Services.
[1023] 5. User Interface:
[1024] The user interface is an interface that allows users to configure the system and check records. A web-based user interface can be built using "React.js."
[1025] To give a specific example, the following system is realized.
[1026] Example 1
[1027] When a visitor enters a store, a camera captures the visitor's video and sends it to a server. The server analyzes the video and identifies the visitor's attributes (age, gender, etc.) and emotions. For example, if the server determines that the visitor is a man in his 40s who is feeling somewhat stressed, it generates a digital human based on that information and displays it on the terminal. The terminal then begins a dialogue with the visitor through the digital human, responding flexibly according to the visitor's emotions. If the visitor wishes to purchase an item and selects QR code payment, the terminal displays the QR code received from the server, and the payment is completed by having the visitor scan it with their smartphone.
[1028] Examples of prompt sentences include:
[1029] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[1030] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1031] Step 1:
[1032] Capture footage of visitors.
[1033] The camera device captures video of visitors in real time and transmits the video data to the server. The input data is raw video data, and the output is video data transmitted to the server. Specifically, the camera device continuously captures video and transmits the data to the server via the network.
[1034] Step 2:
[1035] Analyzing video data and identifying visitor attributes and emotions.
[1036] The server analyzes the transmitted video data and identifies the visitor's attributes (age, gender, etc.). It then uses an emotion engine to analyze the visitor's emotions. The input data is the video data, and the output data is the visitor's attributes and emotional information. Specifically, the server first performs face detection and feature extraction using "OpenCV + Dlib," and then performs emotion analysis on the results using "Microsoft Azure Cognitive Services."
[1037] Step 3:
[1038] Digital human generation.
[1039] The server generates video and audio of a digital human based on the identified attributes and emotions. The input data is the visitor's attributes and emotional information, and the output data is the video and audio data of the generated digital human. Specifically, the video of the digital human is generated using "Unity," and the audio is generated using "voice synthesis technology."
[1040] Step 4:
[1041] Digital human interacting with visitors.
[1042] The terminal displays the digital human received from the server and begins a dialogue with the visitor. The input data is the video and audio data of the digital human, and the output data is the content of the dialogue with the visitor. Specifically, the digital human is displayed on the terminal's display and audio is played back through the audio speaker.
[1043] Step 5:
[1044] Recording and analyzing visitor voices.
[1045] The device records the visitor's voice, analyzes the voice data, and generates an appropriate response. The input data is the visitor's voice data, and the output data is the analyzed voice text and the generated response. Specifically, the device's built-in microphone captures the voice, converts it to text using Google Speech-to-Text, and then the server analyzes the text to generate an appropriate response.
[1046] Step 6:
[1047] Generate and display QR codes.
[1048] When a visitor requests QR code payment, the server generates a dedicated QR code and sends it to the terminal. The terminal then displays it to the visitor. The input data is the visitor's payment preference information, and the output data is the generated QR code. Specifically, the server uses the "Zebra Crossing (ZXing)" library to generate the QR code and display it on the terminal.
[1049] Step 7:
[1050] Recording of conversations and video footage.
[1051] The server records and saves the conversations and video data with visitors. The input data is the conversations and video data, and the output data is the saved information. Specifically, the server uses cloud storage such as AWS S3 to save the data so that it can be viewed later.
[1052] An example of a prompt sentence would be something like:
[1053] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[1054] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1055] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1056] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1057] [Third embodiment]
[1058] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1059] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1060] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1061] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1062] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1063] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1064] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1065] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1066] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1067] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1068] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1069] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1070] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[1071] System Configuration
[1072] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[1073] 2. Server: Receives and analyzes the video data sent from the camera device, identifies visitor attributes, and generates video and audio of a digital human based on those attributes.
[1074] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[1075] 4. User Interface: This is the interface that allows users to configure the system and check records.
[1076] Program processing explanation
[1077] 1. Visitor Recognition
[1078] The terminal transmits the video data from the camera device to the server.
[1079] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1080] 2. Digital Human Generation
[1081] The server generates the appropriate video and audio of the digital human based on the identified attributes.
[1082] For example, the setting is such that elderly people speak in a slow tone, while children speak in a friendly tone.
[1083] 3. Interacting with visitors
[1084] The terminal displays the digital human received from the server and greets the visitor.
[1085] The visitor speaks to the digital human.
[1086] The terminal picks up the visitor's voice and transmits it to the server.
[1087] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[1088] The terminal communicates the generated answer to the visitor via a digital human.
[1089] 4. QR Code Processing
[1090] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[1091] The server generates a QR code and sends it to the device.
[1092] The terminal displays a QR code that visitors can scan with their smartphone.
[1093] 5. Recording and Confirmation
[1094] The server records all conversations and videos.
[1095] Users can later review the recorded conversations and footage through a dedicated interface.
[1096] Specific examples
[1097] Visitor recognition and digital human generation
[1098] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[1099] Interacting with visitors
[1100] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates the appropriate response, "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[1101] QR code processing and recording
[1102] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[1103] The above is an embodiment of the present invention. The present invention automates the process of dealing with visitors, significantly reducing the effort and risk for users.
[1104] The processing flow will be explained below.
[1105] Step 1:
[1106] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[1107] Step 2:
[1108] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[1109] Step 3:
[1110] The server generates the appropriate video and audio of a digital human based on the identified attributes, for example, an elderly person will have a slower voice.
[1111] Step 4:
[1112] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[1113] Step 5:
[1114] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[1115] Step 6:
[1116] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[1117] Step 7:
[1118] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[1119] Step 8:
[1120] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[1121] Step 9:
[1122] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[1123] Step 10:
[1124] The visitor tells the digital human that they would like to process the transaction using a QR code.
[1125] Step 11:
[1126] The server generates a QR code and sends it to the device.
[1127] Step 12:
[1128] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[1129] Step 13:
[1130] The server records all interactions and videos with visitors and stores them in a database.
[1131] Step 14:
[1132] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[1133] Example 1
[1134] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1135] There is a need to build a system that automates visitor interactions and provides appropriate responses based on the visitor's attributes. Conventional visitor interaction systems have difficulty accurately identifying attributes such as the visitor's age and gender and generating a suitable digital human. Furthermore, many systems lack the ability to record and review the content and video of interactions with visitors. Furthermore, there is a lack of a convenient way to make payments and receive items using QR codes.
[1136] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1137] In this invention, the server includes means for analyzing video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, and means for generating a greeting for the digital human using a generative AI model. This makes it possible to generate and display an appropriate digital human according to the visitor attributes, automatically generate a greeting and response according to the attributes, and record and check the content of conversations with visitors and their video.
[1138] A "camera device" is a photographic device for capturing images of visitors, and is installed outside the door, etc.
[1139] "Video data" refers to video information of visitors captured by a camera device.
[1140] "Analyzing video data" means performing processing to identify visitor attribute information (age, gender, belongings, etc.) from the received video data.
[1141] "Visitor attributes" refers to characteristic information about a visitor, such as age, gender, and belongings.
[1142] "Digital Human" means a virtual character, including video and audio, that is generated to interact with visitors.
[1143] An "audio recording device" is a device such as a microphone for capturing the visitor's voice.
[1144] "Voice data" refers to audio information that records the speech of a visitor.
[1145] "Speech recognition technology" is a technology for converting voice data into text and analyzing the content of speech.
[1146] A "generative AI model" is an artificial intelligence model that generates appropriate responses and greetings based on attribute information and speech content.
[1147] A "QR code" is a two-dimensional barcode that visitors use when making payments or receiving items.
[1148] "Recording" refers to the act of saving the content of conversations with visitors and footage.
[1149] "User interface" refers to the interface that allows the user to configure the system and check the recorded dialogue and video.
[1150] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[1151] System Configuration
[1152] 1. Camera device: Installed on the exterior of the door to capture video of visitors. It has the ability to capture high-resolution video, capturing detailed images of visitors' faces and belongings.
[1153] 2. Server: Receives and analyzes the video data sent from the camera. The server incorporates image analysis software (e.g., OpenCV or TensorFlow) and speech recognition technology (e.g., Google Speech-to-Text API). It also generates digital human responses using generative AI models (e.g., OpenAI's GPT-3.5).
[1154] 3. Terminal: Receives data from the server and interacts with visitors. The terminal is equipped with a voice recording device that captures the visitor's voice and sends it to the server.
[1155] 4. User Interface: This is the interface through which users can configure the system, view recorded conversations, and check video footage. It is built as a web application and can be accessed from a browser (e.g., using React or Angular).
[1156] Program processing explanation
[1157] 1. Visitor Recognition:
[1158] The device sends the video data from the camera to the server, which analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1159] 2. Digital Human Generation:
[1160] The server then generates the appropriate video and audio of a digital human based on the identified attributes, and the generated digital human is configured to speak in different tones depending on the visitor's attributes, for example, speaking in a slower tone for elderly people and in a more friendly tone for children.
[1161] 3. Visitor interaction:
[1162] The device displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the device's audio pickup device and sent to the server. The server uses speech recognition technology to analyze the speech and generates an appropriate response using a generative AI model. The device then conveys this response to the visitor via the digital human.
[1163] 4. QR Code Processing:
[1164] When a visitor wishes to make or receive payment using a QR code, they tell the digital human, and the server generates a QR code and sends it to the terminal, which displays the code and the visitor scans it with their smartphone.
[1165] 5. Record and verify:
[1166] The server records all conversations and videos, and users can later review the recorded conversations and videos through a dedicated interface.
[1167] Specific examples
[1168] Visitor recognition and digital human generation
[1169] One day, an elderly visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model designed for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[1170] Interacting with visitors
[1171] The visitor says, "You have a delivery," and the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and generates an appropriate response based on the text: "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[1172] QR code processing and recording
[1173] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[1174] Prompt Sentence Examples
[1175] Here are some example prompts to input to a generative AI model (e.g., OpenAI's GPT-3.5):
[1176] Generate digital human greetings based on visitor attributes. Digital human greetings based on the following attributes:
[1177] Age: Elderly
[1178] Gender: Female
[1179] Possession: Cane
[1180] Please speak in a relaxed, friendly tone for seniors.
[1181] Using this prompt, the system is capable of generating an appropriate digital human greeting for the visitor.
[1182] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1183] Step 1:
[1184] Video capture using camera equipment
[1185] The terminal acquires video data from a camera device installed outside the door.
[1186] Input: Real-time video of the visitor captured by the camera device.
[1187] Output: Captured video data.
[1188] Step 2:
[1189] Video data sent to server
[1190] The terminal compresses the acquired video data, minimizes the amount of data, and transmits it to the server.
[1191] Input: Video data acquired from a camera device.
[1192] Output: Compressed video data sent to the server.
[1193] Step 3:
[1194] Identifying visitor attributes
[1195] The server uses image analysis software (e.g., OpenCV or TensorFlow) to analyze the received video data.
[1196] Through video analysis, the server identifies visitors' attributes such as age, gender, and belongings.
[1197] Input: Compressed video data.
[1198] Output: Visitor demographic information (e.g. age, gender, belongings).
[1199] Step 4:
[1200] Digital Human Generation
[1201] The server uses a generative AI model (e.g., OpenAI's GPT-3.5) to generate a greeting for the digital human based on the visitor's identified attribute information.
[1202] Next, 3D modeling software (e.g., Blender or Unity) is used to generate the video and audio of the digital human.
[1203] Input: Visitor demographic information.
[1204] Output: Digital human video and audio data.
[1205] Step 5:
[1206] Digital human display and greeting
[1207] The terminal displays the video and audio of the digital human received from the server.
[1208] A digital human greets visitors.
[1209] Input: Video and audio data of a digital human.
[1210] Output: A displayed digital human greeting.
[1211] Step 6:
[1212] Interacting with visitors
[1213] The visitor speaks to the digital human.
[1214] The terminal picks up the visitor's speech and sends the audio data to the server.
[1215] Input: Visitor utterance.
[1216] Output: The audio data sent to the server.
[1217] Step 7:
[1218] Voice data analysis and response generation
[1219] The server uses voice recognition technology (e.g., Google Speech-to-Text API) to convert the visitor's speech into text.
[1220] It then uses a generative AI model to generate appropriate responses to what the visitor says.
[1221] Input: Visitor's voice data.
[1222] Output: The recognized text and the response message.
[1223] Step 8:
[1224] Viewing the response
[1225] The terminal conveys the response message received from the server to the visitor through the digital human.
[1226] Input: Digital human's response message.
[1227] Output: Displayed digital human response.
[1228] Step 9:
[1229] QR code processing
[1230] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[1231] The server generates a dedicated QR code and sends it to the device.
[1232] The terminal displays a QR code that visitors can scan with their smartphone.
[1233] Input: Visitor processing request.
[1234] Output: The displayed QR code.
[1235] Step 10:
[1236] Dialogue and video recording
[1237] The server records all conversations and videos.
[1238] Input: Dialogue content and video data.
[1239] Output: Recorded dialogue and video data.
[1240] Step 11:
[1241] Checking the recorded content
[1242] Users can view the recorded conversations and footage through a dedicated user interface.
[1243] Input: Request to access recorded data.
[1244] Output: Displayed recording and video data.
[1245] (Application example 1)
[1246] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1247] Conventional visitor response systems have had several limitations in terms of visitor recognition, interaction, and enhanced security. Specifically, they have difficulty in flexibly interacting with visitors and responding quickly, which increases the workload of users. Furthermore, they have been unable to provide appropriate responses based on visitor attributes or automatically manage records. This has led to problems with systems not functioning effectively, particularly when individualized attention is required for elderly people and children.
[1248] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1249] In this invention, the server includes means for capturing video of nearby visitors with a camera device, means for analyzing the video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, means for interacting with the visitor, means for recording the content of the interaction and the visitor video, means for interacting with the visitor using a smartphone, means for capturing the visitor's audio with the smartphone and analyzing the audio data to generate an appropriate response, and means for generating and displaying a QR code on the smartphone. This enables flexible visitor recognition and appropriate response, reduces user effort, and enhances security.
[1250] A "camera device" is a device for capturing video of visitors in real time.
[1251] "Video data" refers to video information of visitors captured by a camera device.
[1252] "Visitor attributes" refers to information such as the visitor's age, gender, belongings, etc.
[1253] A "digital human" is a video and audio representation of a virtual person generated based on the visitor's attributes.
[1254] "Dialogue" refers to communication between a digital human and a visitor.
[1255] "Recording" refers to saving the content of conversations and footage of visitors.
[1256] A "smartphone" is a portable information terminal that works in conjunction with a camera device and a server and is used to interact with visitors.
[1257] "Audio data" refers to information that records the content of a visitor's speech.
[1258] A "QR code" is a two-dimensional barcode generated for visitors to make payments and other transactions.
[1259] The "server" is a central processing unit that receives data from the camera device, analyzes it, generates digital humans, and records them.
[1260] MODE FOR CARRYING OUT THE INVENTION
[1261] A system for implementing the present invention is configured as follows: First, a camera device captures video of nearby visitors in real time and transmits the video data to a server. Next, the server analyzes the video data and identifies the visitor's attributes. Based on the identified attributes, the server generates video and audio of a digital human and interacts with the visitor via a smartphone.
[1262] Required Hardware and Software
[1263] Camera equipment: Equipment for capturing images of visitors. High-resolution cameras are recommended.
[1264] Server: A central processing unit that analyzes data, generates digital humans, and records data. It incorporates machine learning libraries such as TensorFlow and PyTorch.
[1265] Smartphone: A mobile information device for interacting with visitors. It requires an internet connection, a camera, and a microphone.
[1266] Speech recognition and speech synthesis technologies: The server incorporates speech recognition technologies (such as the Google Speech-to-Text API) and speech synthesis technologies (such as the Google Text-to-Speech API).
[1267] QR code generation software: A library for generating QR codes (e.g., Python's QRCode library).
[1268] Processing flow
[1269] Step 1: Video capture
[1270] The camera captures video of visitors in real time and sends the video data to a server, which analyzes the received video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1271] Step 2: Creating a digital human
[1272] The server generates the appropriate video and audio of the digital human based on the identified attributes. For example, it may be configured to speak in a slower tone for an elderly person, or in a more friendly tone for a child. The data of the generated digital human is then sent to a smartphone.
[1273] Step 3: Dialogue with visitors
[1274] The smartphone displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the smartphone's microphone and sent to the server. The server uses voice recognition technology to analyze the speech and generate an appropriate response. The smartphone then broadcasts the generated response via the digital human.
[1275] Step 4: Processing by QR code
[1276] When a visitor wishes to make a payment or receive payment using a QR code, they notify the digital human. The server generates a dedicated QR code and sends it to the smartphone. The smartphone displays this QR code, and the visitor can scan it to complete the payment process.
[1277] Step 5: Record and verify
[1278] The server records all conversations and video footage, which users can later review through a dedicated interface. This automates interactions with visitors, significantly reducing the user's workload and risk.
[1279] Examples and prompts
[1280] Specific examples
[1281] One day, a visitor approaches the door of a home. A camera installed on the door captures the visitor's video and sends it to a server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people, generates video and audio of the visitor speaking in a slow tone, saying, "Hello. How can I help you?" and sends this video and audio to the smartphone. The smartphone then displays this digital human and greets the visitor.
[1282] Prompt Sentence Examples
[1283] "Generate the appropriate digital human based on the visitor's attributes. For example, set a slower tone for seniors and a more friendly tone for children. Also, generate a QR code."
[1284] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1285] Step 1:
[1286] Input: Video data captured by a camera device
[1287] Processing: The camera device captures video of nearby visitors in real time and collects the video data.
[1288] Output: Collected video data
[1289] Specific operation: The camera captures the moment a visitor approaches the door and saves the video data.
[1290] Step 2:
[1291] Input: Video data obtained from a camera device
[1292] Processing: The server receives the video data and uses video analytics techniques (e.g., facial and object recognition) to identify visitor attributes (e.g., age, gender, belongings, etc.).
[1293] Output: Visitor attribute data identified by analysis
[1294] Specific operation: The server analyzes the video data using TensorFlow and PyTorch, identifies the visitor's face and features, and determines their age and gender.
[1295] Step 3:
[1296] Input: Visitor attribute data
[1297] Processing: The server generates the appropriate video and audio of the digital human based on the visitor's attributes.
[1298] Output: Video and audio data of the generated digital human
[1299] How it works: The server uses the generative AI model to generate a digital human that, for example, greets elderly people in a slower tone. Video and audio data are generated.
[1300] Step 4:
[1301] Input: Digital human video and audio data
[1302] Processing: The server sends the generated digital human data to the smartphone, which displays the digital human to the visitor and greets them.
[1303] Output: Video and audio of a digital human displayed on a smartphone
[1304] Specific operation: The smartphone plays the video of the digital human it receives and greets the user with a greeting such as "Hello. How can I help you?"
[1305] Step 5:
[1306] Input: What the visitor said
[1307] Processing: The device picks up what the visitor is saying and sends the audio data to the server, which uses speech recognition technology to convert it into text data and generate an appropriate response.
[1308] Output: Speech-to-text data and the digital human's response based on it
[1309] Specific operation: The smartphone uses a microphone to pick up the visitor's speech, such as "I have a delivery for you," and the server converts this into text data and generates a response such as "Okay, you can pick it up using the QR code."
[1310] Step 6:
[1311] Input: Visitor's request (e.g., request for pickup via QR code)
[1312] Process: The server generates a QR code and sends it to the smartphone, which displays the generated QR code to the visitor.
[1313] Output: QR code displayed on the smartphone
[1314] How it works: The server generates a QR code using a Python QRCode library or similar, and the smartphone displays the QR code to the visitor.
[1315] Step 7:
[1316] Input: Dialogue content and video data
[1317] Processing: The server records all conversations and video. The user can view the recorded data through a dedicated interface.
[1318] Output: Recorded dialogue and video data
[1319] How it works: The server stores all the interaction data and footage in a database, allowing users to view this data later.
[1320] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1321] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, an emotion engine, and a user interface.
[1322] System Configuration
[1323] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[1324] 2. Server: Receives and analyzes the video and audio data sent from the camera device, identifies the visitor's attributes and emotions, and generates the video and audio of a digital human based on those.
[1325] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[1326] 4. Emotion Engine: An engine that analyzes the visitor's video and audio data and recognizes the visitor's emotions.
[1327] 5. User Interface: This is the interface that allows users to configure the system and check records.
[1328] Program processing explanation
[1329] 1. Visitor Recognition
[1330] The terminal transmits the video data from the camera device to the server.
[1331] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1332] 2. Emotional Recognition
[1333] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[1334] Identify different emotional states of your visitors, such as happy, angry, or anxious.
[1335] 3. Digital Human Generation
[1336] The server generates the appropriate video and audio of the digital human based on the identified attributes and emotions.
[1337] For example, if an elderly visitor is feeling anxious, you might adopt a more friendly and gentle tone.
[1338] 4. Interacting with visitors
[1339] The terminal displays the digital human received from the server and makes initial responses to visitors, such as greetings.
[1340] Visitors speak to the digital human.
[1341] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[1342] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[1343] 5. Regulating responses based on emotions
[1344] Based on the analysis results of the emotion engine, the server adjusts the visual and audio tone of the digital human, for example, responding in a calmer tone to soothe the visitor's anger.
[1345] 6. QR Code Processing
[1346] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[1347] The server generates a QR code and sends it to the device.
[1348] The terminal displays a QR code that visitors can scan with their smartphone.
[1349] 7. Recording and Confirmation
[1350] The server records all conversations and videos.
[1351] The recording also includes visitor sentiment data, which users can review later.
[1352] Users can later review the recorded conversations and footage through a dedicated interface.
[1353] Specific examples
[1354] Visitor recognition and emotion recognition
[1355] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. It also uses an emotion engine to identify that the visitor is feeling anxious. The server then generates a digital human with a gentle tone of voice, specifically for the elderly, to ease their anxiety.
[1356] Interact with visitors and tailor responses based on their emotions
[1357] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates an appropriate response, "Okay, you can pick it up by scanning the QR code." Based on the analysis results of the emotion engine, the response is adjusted to a gentler tone. The device then plays this response on a digital human and conveys it to the visitor.
[1358] QR code processing and recording
[1359] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the device. The device displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server and saved, including emotional data. Users can later review these records through a dedicated interface.
[1360] The above is an embodiment of the present invention. This invention automates the process of interacting with visitors, significantly reducing the user's workload and risk. By combining it with an emotion engine, it becomes possible to respond appropriately to visitors' emotional states.
[1361] The processing flow will be explained below.
[1362] Step 1:
[1363] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[1364] Step 2:
[1365] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[1366] Step 3:
[1367] The server analyzes the visitor's emotions from the video and audio data using an emotion engine, which identifies the visitor's emotional state from their facial expressions and tone of voice.
[1368] Step 4:
[1369] The server generates an appropriate digital human image and voice based on the identified attributes and emotions. For example, a digital human with a gentle tone is generated for an elderly visitor who is anxious.
[1370] Step 5:
[1371] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[1372] Step 6:
[1373] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[1374] Step 7:
[1375] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[1376] Step 8:
[1377] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[1378] Step 9:
[1379] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[1380] Step 10:
[1381] Based on the analysis results of the emotion engine, the server adjusts the video and audio tones of the digital human and responds according to the visitor's emotional state.
[1382] Step 11:
[1383] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[1384] Step 12:
[1385] The visitor tells the digital human that they would like to process the transaction using a QR code.
[1386] Step 13:
[1387] The server generates a QR code and sends it to the device.
[1388] Step 14:
[1389] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[1390] Step 15:
[1391] The server records all interactions and videos with visitors and stores them in a database, including emotional data analyzed by the emotion engine.
[1392] Step 16:
[1393] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[1394] Example 2
[1395] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1396] Automated response systems that interact with visitors are required to accurately recognize the visitor's attributes and emotions and generate an appropriate digital human. It is also important to enable visitors to smoothly make payments and receive payments using QR codes. However, integrating these elements into a single system and making it function smoothly is technically difficult.
[1397] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1398] In this invention, the server includes means for capturing video of nearby visitors using a camera device, means for analyzing the video data and identifying the visitor's attributes, means for analyzing the visitor's emotions from the video and audio data, means for generating video and audio of a digital human based on the visitor's attributes and emotions, means for displaying the digital human and engaging in a dialogue with the visitor, means for capturing the visitor's audio and generating an appropriate response using voice recognition technology, means for adjusting the digital human's response based on the emotion analysis results, means for generating and displaying a QR code when the visitor uses a QR code to make or receive payments, and means for recording the content of the dialogue and the visitor's video. This allows the visitor's attributes and emotions to be accurately recognized, enabling smooth payment and receipt using an appropriate digital human for dialogue and the QR code.
[1399] "Camera Device" means equipment used to capture real-time video of visitors.
[1400] "Video Data" refers to still and video images of Visitors captured by camera devices.
[1401] The "server" is a central processing unit that analyzes video and audio data and identifies the attributes and emotions of visitors.
[1402] "Visitor attributes" refers to characteristic information about visitors, such as age, gender, and belongings.
[1403] The "Emotion Engine" is a software interface for analyzing a visitor's emotional state from video and audio data.
[1404] A "digital human" is an artificial visual and audio character that is generated to interact with visitors.
[1405] "Voice recognition technology" is a technology for converting a visitor's voice into text data.
[1406] A "QR code" is a two-dimensional barcode used by visitors to make payments and receive goods.
[1407] "Dialogue content" refers to the entire content of the conversation between the visitor and the digital human.
[1408] "Recording means" refers to methods and devices for saving the content of interactions and video of visitors.
[1409] In one embodiment of the present invention, the system is composed of a camera device, a server, a terminal, an emotion engine, and a user interface, which allows the system to recognize the attributes and emotions of visitors and generate an appropriate digital human to interact with them.
[1410] System Configuration
[1411] camera equipment
[1412] The camera device is installed outside the door and captures visitors' images in real time. The camera device acquires high-resolution images and transmits them to a terminal.
[1413] server
[1414] The server receives and analyzes the video and audio data sent from the camera. The server has the following functions:
[1415] 1. Video analysis: Using a video analysis library such as OpenCV, we identify visitors' faces and extract attributes such as age, gender, and belongings.
[1416] 2. Emotion analysis: Using Microsoft Azure's emotion recognition API, we analyze visitors' emotions from video and audio data.
[1417] 3. Digital Human Generation: Using tools such as Adobe Character Animator, generate the video and audio of a digital human based on identified attributes and emotions.
[1418] 4. Speech Recognition: Using APIs such as Google Cloud Speech-to-Text, convert the visitor's voice data into text and generate an appropriate response.
[1419] 5. Response adjustment: Adjust the digital human's response based on the emotion analysis results.
[1420] Terminal
[1421] The terminal is the device that receives data from the server and interacts with the visitor. The terminal has the following functions:
[1422] 1. Video display: Displays the digital human received from the server and provides initial greetings and guidance to visitors.
[1423] 2. Microphone: A built-in microphone is included to capture the visitor's speech and send it to the server.
[1424] 3. Display QR code: When a visitor uses a QR code to make or receive a payment, the QR code sent from the server is displayed.
[1425] Emotion Engine
[1426] The emotion engine analyzes the video and audio data of visitors and recognizes their emotions. This engine can identify their emotional state, such as whether they are happy, angry, or anxious.
[1427] User Interface
[1428] Users can use the interface to configure the system and view recordings, which include visitor attributes, emotional state, dialogue, and video.
[1429] Specific operation example
[1430] 1. Recognizing visitors' perceptions and emotions
[1431] One day, a visitor approaches the door of your home. A camera installed on the door captures the visitor's image, and the device sends it to the server.
[1432] The server uses video analysis to determine that the visitor is elderly, and an emotion engine to identify that the visitor is feeling anxious.
[1433] The server is designed for the elderly, generating a digital human with a gentle tone to ease anxiety.
[1434] 2. Interact with visitors and tailor responses based on their emotions
[1435] When a visitor says "I have a delivery," the voice is captured by the device's microphone and sent to the server.
[1436] The server uses speech recognition technology to generate the text "You have a delivery" and then generates the appropriate response "Okay, you can pick it up using the QR code."
[1437] Based on the analysis results of the emotion engine, the response is adjusted to a gentle tone, which the device then plays back to the digital human and conveys to the visitor.
[1438] 3. Processing and recording by QR code
[1439] When a visitor requests to receive their gift using a QR code, the server generates a dedicated QR code and sends it to the device.
[1440] The terminal displays this QR code, and visitors can scan it with their smartphone to complete the pickup.
[1441] All conversations and videos are recorded on a server, including emotional data, and users can view these records through a dedicated interface.
[1442] Prompt Sentence Examples
[1443] 1. Prompt to identify visitor demographics:
[1444] Identify the visitor's age, gender, and belongings from the video data, for example, whether the visitor is carrying a bag.
[1445] 2. Emotion analysis prompts:
[1446] Analyze visitor emotions from this video and audio data. Identify whether your visitors are happy, angry, anxious, etc.
[1447] 3. Prompt when generating a digital human:
[1448] If your visitor is elderly and anxious, generate a digital human that responds in a friendly and gentle tone.
[1449] 4. Speech recognition prompts:
[1450] Convert the visitor's speech into text data. For example, convert the speech "I have a delivery" into text.
[1451] The above is an embodiment of the present invention. The present invention enables accurate recognition of visitor attributes and emotions, and enables smooth payment and receipt using appropriate digital human dialogue and QR codes.
[1452] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1453] Step 1:
[1454] The terminal acquires video data from the camera device and transmits it to the server.
[1455] Input: Real-time video data from a camera device
[1456] Data processing: The device reads the video data from the buffer.
[1457] Output: Video data transmitted to the server through the configured network connection
[1458] Specific operation: A camera device is installed outside the door, and when a visitor approaches, it automatically captures video and the device sends the video data to the server.
[1459] Step 2:
[1460] The server analyzes the received video data and identifies the visitor's attributes.
[1461] Input: Video data sent from the device
[1462] Data processing: Using a video analysis library such as OpenCV, a facial recognition algorithm is applied to extract attributes such as age, gender, and belongings.
[1463] Output: Visitor attribute data (e.g., age: 60, gender: male, belongings: bag)
[1464] What happens: The server performs video analysis to detect the visitor's face and identify their attributes.
[1465] Step 3:
[1466] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[1467] Input: Analyzed video and audio data
[1468] Data processing: Analyze facial expressions and tone of voice using Microsoft Azure's emotion recognition API.
[1469] Output: Visitor emotion data (e.g., anxiety, joy, anger)
[1470] What happens: The server sends data to the emotion engine to identify the visitor's emotion.
[1471] Step 4:
[1472] The server generates a digital human based on the visitor's identified attributes and emotions.
[1473] Input: Visitor demographic and sentiment data
[1474] Data processing: Using Adobe Character Animator or similar software, we generate appropriate digital human images and voices.
[1475] Output: Video and audio data of the generated digital human
[1476] Specific behavior: The server generates a digital human with a friendly and gentle tone that corresponds to the identified attributes (e.g., elderly, anxious state).
[1477] Step 5:
[1478] The terminal displays the digital human received from the server and makes an initial response to the visitor.
[1479] Input: Video and audio data of the digital human sent from the server
[1480] Data processing: Display and play data received by the terminal
[1481] Output: The video and audio of the digital human are displayed and played back to the visitor.
[1482] What it does: The device greets the visitor with the video and audio of a digital human, asking, "Hello, how can I help you?"
[1483] Step 6:
[1484] The visitor speaks to the digital human, and the device captures the audio and sends it to the server.
[1485] Input: Visitor's speech
[1486] Data processing: The device's microphone captures the audio and sends the audio data to the server.
[1487] Output: Audio data sent to the server
[1488] Specific operation: The visitor says "I have a delivery," and the voice is captured by the device and sent to the server.
[1489] Step 7:
[1490] The server analyzes what the visitor says and generates an appropriate response.
[1491] Input: Audio data sent from the device
[1492] Data processing: Converting voice data into text using Google Cloud Speech-to-Text API, etc., and then generating an appropriate response.
[1493] Output: Appropriate response text (e.g., "Okay, you can pick it up with the QR code.")
[1494] What happens: The server converts the voice data into text and generates an appropriate response based on that text.
[1495] Step 8:
[1496] The server adjusts the digital human's response based on the emotion analysis results.
[1497] Input: Response text and sentiment data
[1498] Data processing: Adjust tone and facial expression based on emotion analysis results
[1499] Output: Adjusted response data
[1500] Specific behavior: The server responds with a gentler tone, "Okay, you can pick it up with the QR code."
[1501] Step 9:
[1502] The server generates a QR code, which the device displays.
[1503] Input: Visitor's payment and receipt instructions
[1504] Data processing: Generate a QR code using the Python qrcode library
[1505] Output: Generated QR code data
[1506] Specific operation: When a visitor requests receipt using a QR code, the server generates a QR code and sends it to the terminal, which displays it.
[1507] Step 10:
[1508] The server records all conversations and videos, and the user can check the recorded content.
[1509] Input: Dialogue content, video, emotion data
[1510] Data processing: Recording and saving in a database
[1511] Output: Recorded dialogue and video data
[1512] Specific operation: All conversations and videos are recorded on the server and can be viewed by the user through a dedicated interface.
[1513] (Application example 2)
[1514] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1515] Conventional customer service systems have difficulty recognizing customers' emotions and responding appropriately, limiting the ways to improve customer satisfaction. Furthermore, they lack the ability to respond based on the visitor's attributes and emotions, meaning that even customers who require special attention can only receive a uniform level of service. As a result, customers often become dissatisfied.
[1516] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1517] In this invention, the server includes means for analyzing video data to identify the visitor's attributes and emotions, means for generating video and audio of a digital human based on the visitor's attributes and emotions, and means for interacting with the visitor and adjusting responses based on the emotions, thereby enabling flexible responses according to the customer's emotions.
[1518] "Camera Device" means a device used to capture video of a visitor.
[1519] "Video data" refers to video information captured by a camera device.
[1520] The "server" is a device that analyzes video data, identifies the visitor's attributes and emotions, and generates video and audio of a digital human.
[1521] "Attributes" refer to personal characteristics of visitors, such as their age, gender, and belongings.
[1522] "Emotion" refers to the visitor's mental state, and includes states such as joy, anger, and anxiety.
[1523] A "digital human" is a virtual human image and voice generated to interact with visitors.
[1524] "Dialogue" refers to audio and video communication between the digital human and the visitor.
[1525] "Recording" means saving the content of conversations and footage of visitors.
[1526] A QR code is a type of barcode that visitors can scan with a smartphone or other device to receive information or make payments.
[1527] "Audio Recording Device" means a device used to capture the audio of a visitor.
[1528] "Analysis" is the act of processing received data and extracting information.
[1529] A system for implementing the present invention mainly includes the following components:
[1530] 1. Camera equipment:
[1531] The camera device is installed to capture video of visitors in real time. Any commercially available video camera can be used as the camera device. For example, the Logitech C920 can be used.
[1532] 2. Server:
[1533] The server receives the video data sent from the camera and performs the necessary analysis to identify the visitor's attributes and emotions. Specifically, it uses "OpenCV + Dlib" to analyze the video data, and "Microsoft Azure Cognitive Services" can be used for emotion recognition. The server also uses a combination of "Unity + speech synthesis technology" to generate the video and audio of the digital human.
[1534] 3. Terminal:
[1535] The terminal is a device that receives data sent from the server and enables interaction with visitors. The terminal must be equipped with a voice recording device to capture the visitor's voice and speech recognition technology. For this purpose, "Google Speech-to-Text" can be used. Also, the "Zebra Crossing (ZXing)" library can be used to generate QR codes.
[1536] 4. Emotion Engine:
[1537] The emotion engine is used to identify the emotional state of the visitor by analyzing their video and audio data, and can use emotion recognition technologies such as those from Microsoft Azure Cognitive Services.
[1538] 5. User Interface:
[1539] The user interface is an interface that allows users to configure the system and check records. A web-based user interface can be built using "React.js."
[1540] To give a specific example, the following system is realized.
[1541] Example 1
[1542] When a visitor enters a store, a camera captures the visitor's video and sends it to a server. The server analyzes the video and identifies the visitor's attributes (age, gender, etc.) and emotions. For example, if the server determines that the visitor is a man in his 40s who is feeling somewhat stressed, it generates a digital human based on that information and displays it on the terminal. The terminal then begins a dialogue with the visitor through the digital human, responding flexibly according to the visitor's emotions. If the visitor wishes to purchase an item and selects QR code payment, the terminal displays the QR code received from the server, and the payment is completed by having the visitor scan it with their smartphone.
[1543] Examples of prompt sentences include:
[1544] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[1545] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1546] Step 1:
[1547] Capture footage of visitors.
[1548] The camera device captures video of visitors in real time and transmits the video data to the server. The input data is raw video data, and the output is video data transmitted to the server. Specifically, the camera device continuously captures video and transmits the data to the server via the network.
[1549] Step 2:
[1550] Analyzing video data and identifying visitor attributes and emotions.
[1551] The server analyzes the transmitted video data and identifies the visitor's attributes (age, gender, etc.). It then uses an emotion engine to analyze the visitor's emotions. The input data is the video data, and the output data is the visitor's attributes and emotional information. Specifically, the server first performs face detection and feature extraction using "OpenCV + Dlib," and then performs emotion analysis on the results using "Microsoft Azure Cognitive Services."
[1552] Step 3:
[1553] Digital human generation.
[1554] The server generates video and audio of a digital human based on the identified attributes and emotions. The input data is the visitor's attributes and emotional information, and the output data is the video and audio data of the generated digital human. Specifically, the video of the digital human is generated using "Unity," and the audio is generated using "voice synthesis technology."
[1555] Step 4:
[1556] Digital human interacting with visitors.
[1557] The terminal displays the digital human received from the server and begins a dialogue with the visitor. The input data is the video and audio data of the digital human, and the output data is the content of the dialogue with the visitor. Specifically, the digital human is displayed on the terminal's display and audio is played back through the audio speaker.
[1558] Step 5:
[1559] Recording and analyzing visitor voices.
[1560] The device records the visitor's voice, analyzes the voice data, and generates an appropriate response. The input data is the visitor's voice data, and the output data is the analyzed voice text and the generated response. Specifically, the device's built-in microphone captures the voice, converts it to text using Google Speech-to-Text, and then the server analyzes the text to generate an appropriate response.
[1561] Step 6:
[1562] Generate and display QR codes.
[1563] When a visitor requests QR code payment, the server generates a dedicated QR code and sends it to the terminal. The terminal then displays it to the visitor. The input data is the visitor's payment preference information, and the output data is the generated QR code. Specifically, the server uses the "Zebra Crossing (ZXing)" library to generate the QR code and display it on the terminal.
[1564] Step 7:
[1565] Recording of conversations and video footage.
[1566] The server records and saves the conversations and video data with visitors. The input data is the conversations and video data, and the output data is the saved information. Specifically, the server uses cloud storage such as AWS S3 to save the data so that it can be viewed later.
[1567] An example of a prompt sentence would be something like:
[1568] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[1569] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1570] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1571] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1572] [Fourth embodiment]
[1573] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1574] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1575] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1576] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1577] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1578] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1579] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1580] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1581] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1582] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1583] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1584] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1585] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1586] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[1587] System Configuration
[1588] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[1589] 2. Server: Receives and analyzes the video data sent from the camera device, identifies visitor attributes, and generates video and audio of a digital human based on those attributes.
[1590] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[1591] 4. User Interface: This is the interface that allows users to configure the system and check records.
[1592] Program processing explanation
[1593] 1. Visitor Recognition
[1594] The terminal transmits the video data from the camera device to the server.
[1595] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1596] 2. Digital Human Generation
[1597] The server generates the appropriate video and audio of the digital human based on the identified attributes.
[1598] For example, the setting is such that elderly people speak in a slow tone, while children speak in a friendly tone.
[1599] 3. Interacting with visitors
[1600] The terminal displays the digital human received from the server and greets the visitor.
[1601] The visitor speaks to the digital human.
[1602] The terminal picks up the visitor's voice and transmits it to the server.
[1603] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[1604] The terminal communicates the generated answer to the visitor via a digital human.
[1605] 4. QR Code Processing
[1606] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[1607] The server generates a QR code and sends it to the device.
[1608] The terminal displays a QR code that visitors can scan with their smartphone.
[1609] 5. Recording and Confirmation
[1610] The server records all conversations and videos.
[1611] Users can later review the recorded conversations and footage through a dedicated interface.
[1612] Specific examples
[1613] Visitor recognition and digital human generation
[1614] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[1615] Interacting with visitors
[1616] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates the appropriate response, "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[1617] QR code processing and recording
[1618] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[1619] The above is an embodiment of the present invention. The present invention automates the process of dealing with visitors, significantly reducing the effort and risk for users.
[1620] The processing flow will be explained below.
[1621] Step 1:
[1622] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[1623] Step 2:
[1624] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[1625] Step 3:
[1626] The server generates the appropriate video and audio of a digital human based on the identified attributes, for example, an elderly person will have a slower voice.
[1627] Step 4:
[1628] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[1629] Step 5:
[1630] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[1631] Step 6:
[1632] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[1633] Step 7:
[1634] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[1635] Step 8:
[1636] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[1637] Step 9:
[1638] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[1639] Step 10:
[1640] The visitor tells the digital human that they would like to process the transaction using a QR code.
[1641] Step 11:
[1642] The server generates a QR code and sends it to the device.
[1643] Step 12:
[1644] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[1645] Step 13:
[1646] The server records all interactions and videos with visitors and stores them in a database.
[1647] Step 14:
[1648] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[1649] Example 1
[1650] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1651] There is a need to build a system that automates visitor interactions and provides appropriate responses based on the visitor's attributes. Conventional visitor interaction systems have difficulty accurately identifying attributes such as the visitor's age and gender and generating a suitable digital human. Furthermore, many systems lack the ability to record and review the content and video of interactions with visitors. Furthermore, there is a lack of a convenient way to make payments and receive items using QR codes.
[1652] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1653] In this invention, the server includes means for analyzing video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, and means for generating a greeting for the digital human using a generative AI model. This makes it possible to generate and display an appropriate digital human according to the visitor attributes, automatically generate a greeting and response according to the attributes, and record and check the content of conversations with visitors and their video.
[1654] A "camera device" is a photographic device for capturing images of visitors, and is installed outside the door, etc.
[1655] "Video data" refers to video information of visitors captured by a camera device.
[1656] "Analyzing video data" means performing processing to identify visitor attribute information (age, gender, belongings, etc.) from the received video data.
[1657] "Visitor attributes" refers to characteristic information about a visitor, such as age, gender, and belongings.
[1658] "Digital Human" means a virtual character, including video and audio, that is generated to interact with visitors.
[1659] An "audio recording device" is a device such as a microphone for capturing the visitor's voice.
[1660] "Voice data" refers to audio information that records the speech of a visitor.
[1661] "Speech recognition technology" is a technology for converting voice data into text and analyzing the content of speech.
[1662] A "generative AI model" is an artificial intelligence model that generates appropriate responses and greetings based on attribute information and speech content.
[1663] A "QR code" is a two-dimensional barcode that visitors use when making payments or receiving items.
[1664] "Recording" refers to the act of saving the content of conversations with visitors and footage.
[1665] "User interface" refers to the interface that allows the user to configure the system and check the recorded dialogue and video.
[1666] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, and a user interface.
[1667] System Configuration
[1668] 1. Camera device: Installed on the exterior of the door to capture video of visitors. It has the ability to capture high-resolution video, capturing detailed images of visitors' faces and belongings.
[1669] 2. Server: Receives and analyzes the video data sent from the camera. The server incorporates image analysis software (e.g., OpenCV or TensorFlow) and speech recognition technology (e.g., Google Speech-to-Text API). It also generates digital human responses using generative AI models (e.g., OpenAI's GPT-3.5).
[1670] 3. Terminal: Receives data from the server and interacts with visitors. The terminal is equipped with a voice recording device that captures the visitor's voice and sends it to the server.
[1671] 4. User Interface: This is the interface through which users can configure the system, view recorded conversations, and check video footage. It is built as a web application and can be accessed from a browser (e.g., using React or Angular).
[1672] Program processing explanation
[1673] 1. Visitor Recognition:
[1674] The device sends the video data from the camera to the server, which analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1675] 2. Digital Human Generation:
[1676] The server then generates the appropriate video and audio of a digital human based on the identified attributes, and the generated digital human is configured to speak in different tones depending on the visitor's attributes, for example, speaking in a slower tone for elderly people and in a more friendly tone for children.
[1677] 3. Visitor interaction:
[1678] The device displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the device's audio pickup device and sent to the server. The server uses speech recognition technology to analyze the speech and generates an appropriate response using a generative AI model. The device then conveys this response to the visitor via the digital human.
[1679] 4. QR Code Processing:
[1680] When a visitor wishes to make or receive payment using a QR code, they tell the digital human, and the server generates a QR code and sends it to the terminal, which displays the code and the visitor scans it with their smartphone.
[1681] 5. Record and verify:
[1682] The server records all conversations and videos, and users can later review the recorded conversations and videos through a dedicated interface.
[1683] Specific examples
[1684] Visitor recognition and digital human generation
[1685] One day, an elderly visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model designed for elderly people and generates video and audio that speaks to the visitor in a slow tone, saying, "Hello. How can I help you?" The device displays this digital human and greets the visitor.
[1686] Interacting with visitors
[1687] The visitor says, "You have a delivery," and the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and generates an appropriate response based on the text: "Okay, you can pick it up with the QR code." The device then plays this response back to the digital human and relays it to the visitor.
[1688] QR code processing and recording
[1689] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the terminal. The terminal displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server, and users can view them later through a dedicated interface.
[1690] Prompt Sentence Examples
[1691] Here are some example prompts to input to a generative AI model (e.g., OpenAI's GPT-3.5):
[1692] Generate digital human greetings based on visitor attributes. Digital human greetings based on the following attributes:
[1693] Age: Elderly
[1694] Gender: Female
[1695] Possession: Cane
[1696] Please speak in a relaxed, friendly tone for seniors.
[1697] Using this prompt, the system is capable of generating an appropriate digital human greeting for the visitor.
[1698] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1699] Step 1:
[1700] Video capture using camera equipment
[1701] The terminal acquires video data from a camera device installed outside the door.
[1702] Input: Real-time video of the visitor captured by the camera device.
[1703] Output: Captured video data.
[1704] Step 2:
[1705] Video data sent to server
[1706] The terminal compresses the acquired video data, minimizes the amount of data, and transmits it to the server.
[1707] Input: Video data acquired from a camera device.
[1708] Output: Compressed video data sent to the server.
[1709] Step 3:
[1710] Identifying visitor attributes
[1711] The server uses image analysis software (e.g., OpenCV or TensorFlow) to analyze the received video data.
[1712] Through video analysis, the server identifies visitors' attributes such as age, gender, and belongings.
[1713] Input: Compressed video data.
[1714] Output: Visitor demographic information (e.g. age, gender, belongings).
[1715] Step 4:
[1716] Digital Human Generation
[1717] The server uses a generative AI model (e.g., OpenAI's GPT-3.5) to generate a greeting for the digital human based on the visitor's identified attribute information.
[1718] Next, 3D modeling software (e.g., Blender or Unity) is used to generate the video and audio of the digital human.
[1719] Input: Visitor demographic information.
[1720] Output: Digital human video and audio data.
[1721] Step 5:
[1722] Digital human display and greeting
[1723] The terminal displays the video and audio of the digital human received from the server.
[1724] A digital human greets visitors.
[1725] Input: Video and audio data of a digital human.
[1726] Output: A displayed digital human greeting.
[1727] Step 6:
[1728] Interacting with visitors
[1729] The visitor speaks to the digital human.
[1730] The terminal picks up the visitor's speech and sends the audio data to the server.
[1731] Input: Visitor utterance.
[1732] Output: The audio data sent to the server.
[1733] Step 7:
[1734] Voice data analysis and response generation
[1735] The server uses voice recognition technology (e.g., Google Speech-to-Text API) to convert the visitor's speech into text.
[1736] It then uses a generative AI model to generate appropriate responses to what the visitor says.
[1737] Input: Visitor's voice data.
[1738] Output: The recognized text and the response message.
[1739] Step 8:
[1740] Viewing the response
[1741] The terminal conveys the response message received from the server to the visitor through the digital human.
[1742] Input: Digital human's response message.
[1743] Output: Displayed digital human response.
[1744] Step 9:
[1745] QR code processing
[1746] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[1747] The server generates a dedicated QR code and sends it to the device.
[1748] The terminal displays a QR code that visitors can scan with their smartphone.
[1749] Input: Visitor processing request.
[1750] Output: The displayed QR code.
[1751] Step 10:
[1752] Dialogue and video recording
[1753] The server records all conversations and videos.
[1754] Input: Dialogue content and video data.
[1755] Output: Recorded dialogue and video data.
[1756] Step 11:
[1757] Checking the recorded content
[1758] Users can view the recorded conversations and footage through a dedicated user interface.
[1759] Input: Request to access recorded data.
[1760] Output: Displayed recording and video data.
[1761] (Application example 1)
[1762] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1763] Conventional visitor response systems have had several limitations in terms of visitor recognition, interaction, and enhanced security. Specifically, they have difficulty in flexibly interacting with visitors and responding quickly, which increases the workload of users. Furthermore, they have been unable to provide appropriate responses based on visitor attributes or automatically manage records. This has led to problems with systems not functioning effectively, particularly when individualized attention is required for elderly people and children.
[1764] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1765] In this invention, the server includes means for capturing video of nearby visitors with a camera device, means for analyzing the video data and identifying visitor attributes, means for generating video and audio of a digital human based on the visitor attributes, means for interacting with the visitor, means for recording the content of the interaction and the visitor video, means for interacting with the visitor using a smartphone, means for capturing the visitor's audio with the smartphone and analyzing the audio data to generate an appropriate response, and means for generating and displaying a QR code on the smartphone. This enables flexible visitor recognition and appropriate response, reduces user effort, and enhances security.
[1766] A "camera device" is a device for capturing video of visitors in real time.
[1767] "Video data" refers to video information of visitors captured by a camera device.
[1768] "Visitor attributes" refers to information such as the visitor's age, gender, belongings, etc.
[1769] A "digital human" is a video and audio representation of a virtual person generated based on the visitor's attributes.
[1770] "Dialogue" refers to communication between a digital human and a visitor.
[1771] "Recording" refers to saving the content of conversations and footage of visitors.
[1772] A "smartphone" is a portable information terminal that works in conjunction with a camera device and a server and is used to interact with visitors.
[1773] "Audio data" refers to information that records the content of a visitor's speech.
[1774] A "QR code" is a two-dimensional barcode generated for visitors to make payments and other transactions.
[1775] The "server" is a central processing unit that receives data from the camera device, analyzes it, generates digital humans, and records them.
[1776] MODE FOR CARRYING OUT THE INVENTION
[1777] A system for implementing the present invention is configured as follows: First, a camera device captures video of nearby visitors in real time and transmits the video data to a server. Next, the server analyzes the video data and identifies the visitor's attributes. Based on the identified attributes, the server generates video and audio of a digital human and interacts with the visitor via a smartphone.
[1778] Required Hardware and Software
[1779] Camera equipment: Equipment for capturing images of visitors. High-resolution cameras are recommended.
[1780] Server: A central processing unit that analyzes data, generates digital humans, and records data. It incorporates machine learning libraries such as TensorFlow and PyTorch.
[1781] Smartphone: A mobile information device for interacting with visitors. It requires an internet connection, a camera, and a microphone.
[1782] Speech recognition and speech synthesis technologies: The server incorporates speech recognition technologies (such as the Google Speech-to-Text API) and speech synthesis technologies (such as the Google Text-to-Speech API).
[1783] QR code generation software: A library for generating QR codes (e.g., Python's QRCode library).
[1784] Processing flow
[1785] Step 1: Video capture
[1786] The camera captures video of visitors in real time and sends the video data to a server, which analyzes the received video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1787] Step 2: Creating a digital human
[1788] The server generates the appropriate video and audio of the digital human based on the identified attributes. For example, it may be configured to speak in a slower tone for an elderly person, or in a more friendly tone for a child. The data of the generated digital human is then sent to a smartphone.
[1789] Step 3: Dialogue with visitors
[1790] The smartphone displays the digital human received from the server and greets the visitor. The visitor speaks to the digital human, and the voice is captured by the smartphone's microphone and sent to the server. The server uses voice recognition technology to analyze the speech and generate an appropriate response. The smartphone then broadcasts the generated response via the digital human.
[1791] Step 4: Processing by QR code
[1792] When a visitor wishes to make a payment or receive payment using a QR code, they notify the digital human. The server generates a dedicated QR code and sends it to the smartphone. The smartphone displays this QR code, and the visitor can scan it to complete the payment process.
[1793] Step 5: Record and verify
[1794] The server records all conversations and video footage, which users can later review through a dedicated interface. This automates interactions with visitors, significantly reducing the user's workload and risk.
[1795] Examples and prompts
[1796] Specific examples
[1797] One day, a visitor approaches the door of a home. A camera installed on the door captures the visitor's video and sends it to a server. The server analyzes the video and determines that the visitor is elderly. The server selects a digital human model for elderly people, generates video and audio of the visitor speaking in a slow tone, saying, "Hello. How can I help you?" and sends this video and audio to the smartphone. The smartphone then displays this digital human and greets the visitor.
[1798] Prompt Sentence Examples
[1799] "Generate the appropriate digital human based on the visitor's attributes. For example, set a slower tone for seniors and a more friendly tone for children. Also, generate a QR code."
[1800] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1801] Step 1:
[1802] Input: Video data captured by a camera device
[1803] Processing: The camera device captures video of nearby visitors in real time and collects the video data.
[1804] Output: Collected video data
[1805] Specific operation: The camera captures the moment a visitor approaches the door and saves the video data.
[1806] Step 2:
[1807] Input: Video data obtained from a camera device
[1808] Processing: The server receives the video data and uses video analytics techniques (e.g., facial and object recognition) to identify visitor attributes (e.g., age, gender, belongings, etc.).
[1809] Output: Visitor attribute data identified by analysis
[1810] Specific operation: The server analyzes the video data using TensorFlow and PyTorch, identifies the visitor's face and features, and determines their age and gender.
[1811] Step 3:
[1812] Input: Visitor attribute data
[1813] Processing: The server generates the appropriate video and audio of the digital human based on the visitor's attributes.
[1814] Output: Video and audio data of the generated digital human
[1815] How it works: The server uses the generative AI model to generate a digital human that, for example, greets elderly people in a slower tone. Video and audio data are generated.
[1816] Step 4:
[1817] Input: Digital human video and audio data
[1818] Processing: The server sends the generated digital human data to the smartphone, which displays the digital human to the visitor and greets them.
[1819] Output: Video and audio of a digital human displayed on a smartphone
[1820] Specific operation: The smartphone plays the video of the digital human it receives and greets the user with a greeting such as "Hello. How can I help you?"
[1821] Step 5:
[1822] Input: What the visitor said
[1823] Processing: The device picks up what the visitor is saying and sends the audio data to the server, which uses speech recognition technology to convert it into text data and generate an appropriate response.
[1824] Output: Speech-to-text data and the digital human's response based on it
[1825] Specific operation: The smartphone uses a microphone to pick up the visitor's speech, such as "I have a delivery for you," and the server converts this into text data and generates a response such as "Okay, you can pick it up using the QR code."
[1826] Step 6:
[1827] Input: Visitor's request (e.g., request for pickup via QR code)
[1828] Process: The server generates a QR code and sends it to the smartphone, which displays the generated QR code to the visitor.
[1829] Output: QR code displayed on the smartphone
[1830] How it works: The server generates a QR code using a Python QRCode library or similar, and the smartphone displays the QR code to the visitor.
[1831] Step 7:
[1832] Input: Dialogue content and video data
[1833] Processing: The server records all conversations and video. The user can view the recorded data through a dedicated interface.
[1834] Output: Recorded dialogue and video data
[1835] How it works: The server stores all the interaction data and footage in a database, allowing users to view this data later.
[1836] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1837] As an embodiment of the present invention, the system is configured as follows: The system is configured from a camera device, a server, a terminal, an emotion engine, and a user interface.
[1838] System Configuration
[1839] 1. Camera device: This device is installed on the outside of the door and captures video of visitors in real time.
[1840] 2. Server: Receives and analyzes the video and audio data sent from the camera device, identifies the visitor's attributes and emotions, and generates the video and audio of a digital human based on those.
[1841] 3. Terminal: A device that receives data from the server and interacts with visitors. The terminal is equipped with voice recognition and voice synthesis technologies.
[1842] 4. Emotion Engine: An engine that analyzes the visitor's video and audio data and recognizes the visitor's emotions.
[1843] 5. User Interface: This is the interface that allows users to configure the system and check records.
[1844] Program processing explanation
[1845] 1. Visitor Recognition
[1846] The terminal transmits the video data from the camera device to the server.
[1847] The server analyzes the video data and identifies the visitor's attributes (age, gender, belongings, etc.).
[1848] 2. Emotional Recognition
[1849] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[1850] Identify different emotional states of your visitors, such as happy, angry, or anxious.
[1851] 3. Digital Human Generation
[1852] The server generates the appropriate video and audio of the digital human based on the identified attributes and emotions.
[1853] For example, if an elderly visitor is feeling anxious, you might adopt a more friendly and gentle tone.
[1854] 4. Interacting with visitors
[1855] The terminal displays the digital human received from the server and makes initial responses to visitors, such as greetings.
[1856] Visitors speak to the digital human.
[1857] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[1858] The server uses voice recognition technology to analyze the speech and generate an appropriate response.
[1859] 5. Regulating responses based on emotions
[1860] Based on the analysis results of the emotion engine, the server adjusts the visual and audio tone of the digital human, for example, responding in a calmer tone to soothe the visitor's anger.
[1861] 6. QR Code Processing
[1862] If a visitor wishes to make or receive payment using a QR code, they will communicate this to the digital human.
[1863] The server generates a QR code and sends it to the device.
[1864] The terminal displays a QR code that visitors can scan with their smartphone.
[1865] 7. Recording and Confirmation
[1866] The server records all conversations and videos.
[1867] The recording also includes visitor sentiment data, which users can review later.
[1868] Users can later review the recorded conversations and footage through a dedicated interface.
[1869] Specific examples
[1870] Visitor recognition and emotion recognition
[1871] One day, a visitor approaches the door of the home. A camera installed on the door captures the visitor's video, which the device sends to the server. The server analyzes the video and determines that the visitor is elderly. It also uses an emotion engine to identify that the visitor is feeling anxious. The server then generates a digital human with a gentle tone of voice, specifically for the elderly, to ease their anxiety.
[1872] Interact with visitors and tailor responses based on their emotions
[1873] When a visitor says, "You have a delivery," the voice is captured by the device's microphone and sent to the server. The server uses speech recognition technology to generate the text "You have a delivery," and then generates an appropriate response, "Okay, you can pick it up by scanning the QR code." Based on the analysis results of the emotion engine, the response is adjusted to a gentler tone. The device then plays this response on a digital human and conveys it to the visitor.
[1874] QR code processing and recording
[1875] When a visitor requests to receive their item using a QR code, the server generates a dedicated QR code and sends it to the device. The device displays the QR code, and the visitor scans it with their smartphone to complete the delivery. All conversations and video are recorded on the server and saved, including emotional data. Users can later review these records through a dedicated interface.
[1876] The above is an embodiment of the present invention. This invention automates the process of interacting with visitors, significantly reducing the user's workload and risk. By combining it with an emotion engine, it becomes possible to respond appropriately to visitors' emotional states.
[1877] The processing flow will be explained below.
[1878] Step 1:
[1879] The terminal uses a camera device to capture video of the visitor in real time and transmits the data to a server.
[1880] Step 2:
[1881] The server analyzes the received video data and performs video analysis to identify the visitor's attributes (e.g., age, gender, belongings).
[1882] Step 3:
[1883] The server analyzes the visitor's emotions from the video and audio data using an emotion engine, which identifies the visitor's emotional state from their facial expressions and tone of voice.
[1884] Step 4:
[1885] The server generates an appropriate digital human image and voice based on the identified attributes and emotions. For example, a digital human with a gentle tone is generated for an elderly visitor who is anxious.
[1886] Step 5:
[1887] The terminal displays the video and audio of the digital human received from the server and makes initial responses to visitors, such as greetings.
[1888] Step 6:
[1889] The visitor speaks to the digital human, for example, saying, "I have a delivery for you."
[1890] Step 7:
[1891] The device uses a microphone to capture the visitor's voice and transmits the data to the server.
[1892] Step 8:
[1893] The server uses voice recognition technology to convert the visitor's speech into text and analyzes the content.
[1894] Step 9:
[1895] The server generates an appropriate response based on the analysis results and generates a voice using the speech synthesis engine again, for example, "I understand. You can receive it by scanning the QR code."
[1896] Step 10:
[1897] Based on the analysis results of the emotion engine, the server adjusts the video and audio tones of the digital human and responds according to the visitor's emotional state.
[1898] Step 11:
[1899] The terminal displays the response received from the server to the digital human and plays it back to the visitor.
[1900] Step 12:
[1901] The visitor tells the digital human that they would like to process the transaction using a QR code.
[1902] Step 13:
[1903] The server generates a QR code and sends it to the device.
[1904] Step 14:
[1905] The terminal will display the generated QR code on the screen for visitors to scan with their smartphones.
[1906] Step 15:
[1907] The server records all interactions and videos with visitors and stores them in a database, including emotional data analyzed by the emotion engine.
[1908] Step 16:
[1909] Users can check the recorded conversations and videos as needed through a dedicated user interface.
[1910] Example 2
[1911] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1912] Automated response systems that interact with visitors are required to accurately recognize the visitor's attributes and emotions and generate an appropriate digital human. It is also important to enable visitors to smoothly make payments and receive payments using QR codes. However, integrating these elements into a single system and making it function smoothly is technically difficult.
[1913] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1914] In this invention, the server includes means for capturing video of nearby visitors using a camera device, means for analyzing the video data and identifying the visitor's attributes, means for analyzing the visitor's emotions from the video and audio data, means for generating video and audio of a digital human based on the visitor's attributes and emotions, means for displaying the digital human and engaging in a dialogue with the visitor, means for capturing the visitor's audio and generating an appropriate response using voice recognition technology, means for adjusting the digital human's response based on the emotion analysis results, means for generating and displaying a QR code when the visitor uses a QR code to make or receive payments, and means for recording the content of the dialogue and the visitor's video. This allows the visitor's attributes and emotions to be accurately recognized, enabling smooth payment and receipt using an appropriate digital human for dialogue and the QR code.
[1915] "Camera Device" means equipment used to capture real-time video of visitors.
[1916] "Video Data" refers to still and video images of Visitors captured by camera devices.
[1917] The "server" is a central processing unit that analyzes video and audio data and identifies the attributes and emotions of visitors.
[1918] "Visitor attributes" refers to characteristic information about visitors, such as age, gender, and belongings.
[1919] The "Emotion Engine" is a software interface for analyzing a visitor's emotional state from video and audio data.
[1920] A "digital human" is an artificial visual and audio character that is generated to interact with visitors.
[1921] "Voice recognition technology" is a technology for converting a visitor's voice into text data.
[1922] A "QR code" is a two-dimensional barcode used by visitors to make payments and receive goods.
[1923] "Dialogue content" refers to the entire content of the conversation between the visitor and the digital human.
[1924] "Recording means" refers to methods and devices for saving the content of interactions and video of visitors.
[1925] In one embodiment of the present invention, the system is composed of a camera device, a server, a terminal, an emotion engine, and a user interface, which allows the system to recognize the attributes and emotions of visitors and generate an appropriate digital human to interact with them.
[1926] System Configuration
[1927] camera equipment
[1928] The camera device is installed outside the door and captures visitors' images in real time. The camera device acquires high-resolution images and transmits them to a terminal.
[1929] server
[1930] The server receives and analyzes the video and audio data sent from the camera. The server has the following functions:
[1931] 1. Video analysis: Using a video analysis library such as OpenCV, we identify visitors' faces and extract attributes such as age, gender, and belongings.
[1932] 2. Emotion analysis: Using Microsoft Azure's emotion recognition API, we analyze visitors' emotions from video and audio data.
[1933] 3. Digital Human Generation: Using tools such as Adobe Character Animator, generate the video and audio of a digital human based on identified attributes and emotions.
[1934] 4. Speech Recognition: Using APIs such as Google Cloud Speech-to-Text, convert the visitor's voice data into text and generate an appropriate response.
[1935] 5. Response adjustment: Adjust the digital human's response based on the emotion analysis results.
[1936] Terminal
[1937] The terminal is the device that receives data from the server and interacts with the visitor. The terminal has the following functions:
[1938] 1. Video display: Displays the digital human received from the server and provides initial greetings and guidance to visitors.
[1939] 2. Microphone: A built-in microphone is included to capture the visitor's speech and send it to the server.
[1940] 3. Display QR code: When a visitor uses a QR code to make or receive a payment, the QR code sent from the server is displayed.
[1941] Emotion Engine
[1942] The emotion engine analyzes the video and audio data of visitors and recognizes their emotions. This engine can identify their emotional state, such as whether they are happy, angry, or anxious.
[1943] User Interface
[1944] Users can use the interface to configure the system and view recordings, which include visitor attributes, emotional state, dialogue, and video.
[1945] Specific operation example
[1946] 1. Recognizing visitors' perceptions and emotions
[1947] One day, a visitor approaches the door of your home. A camera installed on the door captures the visitor's image, and the device sends it to the server.
[1948] The server uses video analysis to determine that the visitor is elderly, and an emotion engine to identify that the visitor is feeling anxious.
[1949] The server is designed for the elderly, generating a digital human with a gentle tone to ease anxiety.
[1950] 2. Interact with visitors and tailor responses based on their emotions
[1951] When a visitor says "I have a delivery," the voice is captured by the device's microphone and sent to the server.
[1952] The server uses speech recognition technology to generate the text "You have a delivery" and then generates the appropriate response "Okay, you can pick it up using the QR code."
[1953] Based on the analysis results of the emotion engine, the response is adjusted to a gentle tone, which the device then plays back to the digital human and conveys to the visitor.
[1954] 3. Processing and recording by QR code
[1955] When a visitor requests to receive their gift using a QR code, the server generates a dedicated QR code and sends it to the device.
[1956] The terminal displays this QR code, and visitors can scan it with their smartphone to complete the pickup.
[1957] All conversations and videos are recorded on a server, including emotional data, and users can view these records through a dedicated interface.
[1958] Prompt Sentence Examples
[1959] 1. Prompt to identify visitor demographics:
[1960] Identify the visitor's age, gender, and belongings from the video data, for example, whether the visitor is carrying a bag.
[1961] 2. Emotion analysis prompts:
[1962] Analyze visitor emotions from this video and audio data. Identify whether your visitors are happy, angry, anxious, etc.
[1963] 3. Prompt when generating a digital human:
[1964] If your visitor is elderly and anxious, generate a digital human that responds in a friendly and gentle tone.
[1965] 4. Speech recognition prompts:
[1966] Convert the visitor's speech into text data. For example, convert the speech "I have a delivery" into text.
[1967] The above is an embodiment of the present invention. The present invention enables accurate recognition of visitor attributes and emotions, and enables smooth payment and receipt using appropriate digital human dialogue and QR codes.
[1968] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1969] Step 1:
[1970] The terminal acquires video data from the camera device and transmits it to the server.
[1971] Input: Real-time video data from a camera device
[1972] Data processing: The device reads the video data from the buffer.
[1973] Output: Video data transmitted to the server through the configured network connection
[1974] Specific operation: A camera device is installed outside the door, and when a visitor approaches, it automatically captures video and the device sends the video data to the server.
[1975] Step 2:
[1976] The server analyzes the received video data and identifies the visitor's attributes.
[1977] Input: Video data sent from the device
[1978] Data processing: Using a video analysis library such as OpenCV, a facial recognition algorithm is applied to extract attributes such as age, gender, and belongings.
[1979] Output: Visitor attribute data (e.g., age: 60, gender: male, belongings: bag)
[1980] What happens: The server performs video analysis to detect the visitor's face and identify their attributes.
[1981] Step 3:
[1982] The server uses an emotion engine to analyze the visitor's emotions from the video and audio data.
[1983] Input: Analyzed video and audio data
[1984] Data processing: Analyze facial expressions and tone of voice using Microsoft Azure's emotion recognition API.
[1985] Output: Visitor emotion data (e.g., anxiety, joy, anger)
[1986] What happens: The server sends data to the emotion engine to identify the visitor's emotion.
[1987] Step 4:
[1988] The server generates a digital human based on the visitor's identified attributes and emotions.
[1989] Input: Visitor demographic and sentiment data
[1990] Data processing: Using Adobe Character Animator or similar software, we generate appropriate digital human images and voices.
[1991] Output: Video and audio data of the generated digital human
[1992] Specific behavior: The server generates a digital human with a friendly and gentle tone that corresponds to the identified attributes (e.g., elderly, anxious state).
[1993] Step 5:
[1994] The terminal displays the digital human received from the server and makes an initial response to the visitor.
[1995] Input: Video and audio data of the digital human sent from the server
[1996] Data processing: Display and play data received by the terminal
[1997] Output: The video and audio of the digital human are displayed and played back to the visitor.
[1998] What it does: The device greets the visitor with the video and audio of a digital human, asking, "Hello, how can I help you?"
[1999] Step 6:
[2000] The visitor speaks to the digital human, and the device captures the audio and sends it to the server.
[2001] Input: Visitor's speech
[2002] Data processing: The device's microphone captures the audio and sends the audio data to the server.
[2003] Output: Audio data sent to the server
[2004] Specific operation: The visitor says "I have a delivery," and the voice is captured by the device and sent to the server.
[2005] Step 7:
[2006] The server analyzes what the visitor says and generates an appropriate response.
[2007] Input: Audio data sent from the device
[2008] Data processing: Converting voice data into text using Google Cloud Speech-to-Text API, etc., and then generating an appropriate response.
[2009] Output: Appropriate response text (e.g., "Okay, you can pick it up with the QR code.")
[2010] What happens: The server converts the voice data into text and generates an appropriate response based on that text.
[2011] Step 8:
[2012] The server adjusts the digital human's response based on the emotion analysis results.
[2013] Input: Response text and sentiment data
[2014] Data processing: Adjust tone and facial expression based on emotion analysis results
[2015] Output: Adjusted response data
[2016] Specific behavior: The server responds with a gentler tone, "Okay, you can pick it up with the QR code."
[2017] Step 9:
[2018] The server generates a QR code, which the device displays.
[2019] Input: Visitor's payment and receipt instructions
[2020] Data processing: Generate a QR code using the Python qrcode library
[2021] Output: Generated QR code data
[2022] Specific operation: When a visitor requests receipt using a QR code, the server generates a QR code and sends it to the terminal, which displays it.
[2023] Step 10:
[2024] The server records all conversations and videos, and the user can check the recorded content.
[2025] Input: Dialogue content, video, emotion data
[2026] Data processing: Recording and saving in a database
[2027] Output: Recorded dialogue and video data
[2028] Specific operation: All conversations and videos are recorded on the server and can be viewed by the user through a dedicated interface.
[2029] (Application example 2)
[2030] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2031] Conventional customer service systems have difficulty recognizing customers' emotions and responding appropriately, limiting the ways to improve customer satisfaction. Furthermore, they lack the ability to respond based on the visitor's attributes and emotions, meaning that even customers who require special attention can only receive a uniform level of service. As a result, customers often become dissatisfied.
[2032] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2033] In this invention, the server includes means for analyzing video data to identify the visitor's attributes and emotions, means for generating video and audio of a digital human based on the visitor's attributes and emotions, and means for interacting with the visitor and adjusting responses based on the emotions, thereby enabling flexible responses according to the customer's emotions.
[2034] "Camera Device" means a device used to capture video of a visitor.
[2035] "Video data" refers to video information captured by a camera device.
[2036] The "server" is a device that analyzes video data, identifies the visitor's attributes and emotions, and generates video and audio of a digital human.
[2037] "Attributes" refer to personal characteristics of visitors, such as their age, gender, and belongings.
[2038] "Emotion" refers to the visitor's mental state, and includes states such as joy, anger, and anxiety.
[2039] A "digital human" is a virtual human image and voice generated to interact with visitors.
[2040] "Dialogue" refers to audio and video communication between the digital human and the visitor.
[2041] "Recording" means saving the content of conversations and footage of visitors.
[2042] A QR code is a type of barcode that visitors can scan with a smartphone or other device to receive information or make payments.
[2043] "Audio Recording Device" means a device used to capture the audio of a visitor.
[2044] "Analysis" is the act of processing received data and extracting information.
[2045] A system for implementing the present invention mainly includes the following components:
[2046] 1. Camera equipment:
[2047] The camera device is installed to capture video of visitors in real time. Any commercially available video camera can be used as the camera device. For example, the Logitech C920 can be used.
[2048] 2. Server:
[2049] The server receives the video data sent from the camera and performs the necessary analysis to identify the visitor's attributes and emotions. Specifically, it uses "OpenCV + Dlib" to analyze the video data, and "Microsoft Azure Cognitive Services" can be used for emotion recognition. The server also uses a combination of "Unity + speech synthesis technology" to generate the video and audio of the digital human.
[2050] 3. Terminal:
[2051] The terminal is a device that receives data sent from the server and enables interaction with visitors. The terminal must be equipped with a voice recording device to capture the visitor's voice and speech recognition technology. For this purpose, "Google Speech-to-Text" can be used. Also, the "Zebra Crossing (ZXing)" library can be used to generate QR codes.
[2052] 4. Emotion Engine:
[2053] The emotion engine is used to identify the emotional state of the visitor by analyzing their video and audio data, and can use emotion recognition technologies such as those from Microsoft Azure Cognitive Services.
[2054] 5. User Interface:
[2055] The user interface is an interface that allows users to configure the system and check records. A web-based user interface can be built using "React.js."
[2056] To give a specific example, the following system is realized.
[2057] Example 1
[2058] When a visitor enters a store, a camera captures the visitor's video and sends it to a server. The server analyzes the video and identifies the visitor's attributes (age, gender, etc.) and emotions. For example, if the server determines that the visitor is a man in his 40s who is feeling somewhat stressed, it generates a digital human based on that information and displays it on the terminal. The terminal then begins a dialogue with the visitor through the digital human, responding flexibly according to the visitor's emotions. If the visitor wishes to purchase an item and selects QR code payment, the terminal displays the QR code received from the server, and the payment is completed by having the visitor scan it with their smartphone.
[2059] Examples of prompt sentences include:
[2060] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[2061] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2062] Step 1:
[2063] Capture footage of visitors.
[2064] The camera device captures video of visitors in real time and transmits the video data to the server. The input data is raw video data, and the output is video data transmitted to the server. Specifically, the camera device continuously captures video and transmits the data to the server via the network.
[2065] Step 2:
[2066] Analyzing video data and identifying visitor attributes and emotions.
[2067] The server analyzes the transmitted video data and identifies the visitor's attributes (age, gender, etc.). It then uses an emotion engine to analyze the visitor's emotions. The input data is the video data, and the output data is the visitor's attributes and emotional information. Specifically, the server first performs face detection and feature extraction using "OpenCV + Dlib," and then performs emotion analysis on the results using "Microsoft Azure Cognitive Services."
[2068] Step 3:
[2069] Digital human generation.
[2070] The server generates video and audio of a digital human based on the identified attributes and emotions. The input data is the visitor's attributes and emotional information, and the output data is the video and audio data of the generated digital human. Specifically, the video of the digital human is generated using "Unity," and the audio is generated using "voice synthesis technology."
[2071] Step 4:
[2072] Digital human interacting with visitors.
[2073] The terminal displays the digital human received from the server and begins a dialogue with the visitor. The input data is the video and audio data of the digital human, and the output data is the content of the dialogue with the visitor. Specifically, the digital human is displayed on the terminal's display and audio is played back through the audio speaker.
[2074] Step 5:
[2075] Recording and analyzing visitor voices.
[2076] The device records the visitor's voice, analyzes the voice data, and generates an appropriate response. The input data is the visitor's voice data, and the output data is the analyzed voice text and the generated response. Specifically, the device's built-in microphone captures the voice, converts it to text using Google Speech-to-Text, and then the server analyzes the text to generate an appropriate response.
[2077] Step 6:
[2078] Generate and display QR codes.
[2079] When a visitor requests QR code payment, the server generates a dedicated QR code and sends it to the terminal. The terminal then displays it to the visitor. The input data is the visitor's payment preference information, and the output data is the generated QR code. Specifically, the server uses the "Zebra Crossing (ZXing)" library to generate the QR code and display it on the terminal.
[2080] Step 7:
[2081] Recording of conversations and video footage.
[2082] The server records and saves the conversations and video data with visitors. The input data is the conversations and video data, and the output data is the saved information. Specifically, the server uses cloud storage such as AWS S3 to save the data so that it can be viewed later.
[2083] An example of a prompt sentence would be something like:
[2084] "This customer is a man in his 40s who seems to be feeling a bit stressed. Based on this information, please use a digital human to generate advice on how to best serve him."
[2085] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2086] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2087] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2088] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2089] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2090] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2091] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2092] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2093] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2094] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2095] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2096] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2097] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2098] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2099] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2100] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2101] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2102] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2103] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2104] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2105] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2106] The following is further disclosed regarding the above embodiment.
[2107] (Claim 1)
[2108] means for capturing video of nearby visitors with a camera device;
[2109] A means for analyzing video data and identifying visitor attributes;
[2110] means for generating video and audio of a digital human based on visitor attributes;
[2111] a means of interacting with visitors;
[2112] a means for recording the content of the conversation and video of the visitor;
[2113] means for generating and displaying a QR code;
[2114] A system including:
[2115] (Claim 2)
[2116] 10. The system of claim 1, wherein the camera device is located on the exterior of the door.
[2117] (Claim 3)
[2118] 10. The system of claim 1, further comprising means for capturing the visitor's voice using an audio recording device and generating an appropriate response by analyzing the voice data.
[2119] "Example 1"
[2120] (Claim 1)
[2121] means for capturing video of nearby visitors with a camera device;
[2122] A means for analyzing video data and identifying visitor attributes;
[2123] means for generating video and audio of a digital human based on visitor attributes;
[2124] a means of interacting with visitors;
[2125] means for capturing the visitor's voice using a voice recording device and generating an appropriate response by analyzing the voice data;
[2126] a means for recording the content of the conversation and video of the visitor;
[2127] means for generating and displaying a QR code;
[2128] A means for users to review the recorded conversations and footage;
[2129] A system including:
[2130] (Claim 2)
[2131] 10. The system of claim 1, wherein the camera device is located on the exterior of the door.
[2132] (Claim 3)
[2133] 10. The system of claim 1, further comprising: means for generating a greeting for the digital human using a generative AI model based on attributes.
[2134] "Application Example 1"
[2135] (Claim 1)
[2136] means for capturing video of nearby visitors with a camera device;
[2137] A means for analyzing video data and identifying visitor attributes;
[2138] means for generating video and audio of a digital human based on visitor attributes;
[2139] a means of interacting with visitors;
[2140] a means for recording the content of the conversation and video of the visitor;
[2141] A means for interacting with visitors using a smartphone;
[2142] A means of capturing the visitor's voice using a smartphone and analyzing the voice data to generate an appropriate response;
[2143] A means for generating and displaying a QR code on a smartphone;
[2144] A system including:
[2145] (Claim 2)
[2146] 10. The system of claim 1, wherein the camera device is located on the exterior of the door.
[2147] (Claim 3)
[2148] 10. The system of claim 1, further comprising means for capturing the visitor's voice using an audio recording device and generating an appropriate response by analyzing the voice data.
[2149] "Example 2: Combining Emotion Engines"
[2150] (Claim 1)
[2151] means for capturing video of nearby visitors with a camera device;
[2152] A means for analyzing video data and identifying visitor attributes;
[2153] A means for analyzing visitor emotions from video and audio data;
[2154] means for generating video and audio of a digital human based on visitor attributes and emotions;
[2155] a means for displaying a digital human and interacting with visitors;
[2156] means for capturing the visitor's voice and generating an appropriate response using voice recognition technology;
[2157] means for adjusting the response of the digital human based on the emotion analysis results;
[2158] A means for generating and displaying a QR code when a visitor uses the QR code to make or receive a payment;
[2159] a means for recording the content of the conversation and video of the visitor;
[2160] A system including:
[2161] (Claim 2)
[2162] 10. The system of claim 1, wherein the camera device is located on the exterior of the door.
[2163] (Claim 3)
[2164] 10. The system of claim 1, further comprising means for capturing the visitor's voice using an audio recording device and generating an appropriate response by analyzing the voice data.
[2165] "Application example 2 when combining emotion engines"
[2166] (Claim 1)
[2167] means for capturing video of nearby visitors with a camera device;
[2168] A means for analyzing the video data and identifying visitor attributes and emotions;
[2169] means for generating video and audio of a digital human based on visitor attributes and emotions;
[2170] A means of interacting with visitors and tailoring responses based on their emotions;
[2171] a means for recording the content of the conversation and video of the visitor;
[2172] means for generating and displaying a QR code;
[2173] A system including:
[2174] (Claim 2)
[2175] 2. The system of claim 1, wherein the camera device is installed outside the door or at the entrance of the store.
[2176] (Claim 3)
[2177] 10. The system of claim 1, further comprising means for capturing the visitor's voice using an audio recording device, generating an appropriate response by analyzing the voice data, and further adjusting the response using an emotion engine. [Explanation of symbols]
[2178] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for capturing video of nearby visitors with a camera device; A means for analyzing video data and identifying visitor attributes; means for generating video and audio of a digital human based on visitor attributes; a means of interacting with visitors; a means for recording the content of the conversation and video of the visitor; means for generating and displaying a QR code; A system including:
2. 2. The system of claim 1, wherein the camera device is located on the exterior of the door.
3. 10. The system of claim 1, further comprising means for capturing the visitor's voice using a voice recording device and generating an appropriate response by analyzing the voice data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A