System

The system addresses the challenge of identifying visitors and providing personalized responses by capturing images, transmitting data for identification, and generating voice messages, enhancing both security and hospitality.

JP2026022288APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123805
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing systems fail to provide effective security measures and hospitality when guests return home, as they lack the ability to identify visitors and respond appropriately, leading to user dissatisfaction and anxiety while the home is empty.

Method used

A system that captures visitor images using a camera, transmits data to a server for identification via image recognition, and generates voice messages based on the identification results, distinguishing between registered and unregistered users to provide personalized greetings or warnings.

Benefits of technology

The system enhances security and hospitality by accurately identifying visitors and providing tailored voice messages, reducing user anxiety and improving the overall home experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022288000001_ABST
    Figure 2026022288000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for capturing an image of a visitor using a camera installed near an entrance; means for transmitting the captured image data to a server; means for identifying the visitor using image recognition technology on the server; means for determining whether the visitor is a registered user on the server; and means for generating and outputting an audio message based on a determination result from the server.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] There is a lack of effective systems that simultaneously provide security measures and improve hospitality when guests return home in modern homes. In particular, there is a need for a way to check who has visited while the home is empty and a way to reduce anxiety when the home is empty. It is also expected that automated greetings will provide a user-friendly hospitality experience when guests return home. Conventional security cameras and surveillance systems have limited ability to identify visitors and respond appropriately, making it difficult to achieve sufficient user satisfaction. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for capturing images of visitors using a camera installed near the entrance, a means for transmitting the captured image data to a server, a means for identifying the visitor using image recognition technology on the server, a means for determining whether the visitor is a registered user on the server, and a means for generating and outputting a voice message based on the determination result from the server. This allows registered users to receive a voice message saying "Welcome back," while warning unregistered visitors that "Welcome. This is being recorded," thereby achieving both security measures and hospitality. Furthermore, by providing a function for checking the status of the home from outside, the system reduces anxiety while the user is away and supports daily life.

[0006] A "camera" is a device installed near the entrance to capture facial images of visitors.

[0007] The "server" is a processing device that receives the captured image data, identifies the visitor using image recognition technology, and generates a voice message based on the identification result.

[0008] "Image data" is data including facial image information of visitors captured by a camera.

[0009] "Image recognition technology" refers to algorithms and software technologies that identify visitors' faces from captured image data.

[0010] A "voice message" is a voice guidance or warning message that is generated based on the server's judgment results and output from the terminal.

[0011] A "registered user" refers to a specific visitor whose facial image data has been pre-registered in the system.

[0012] An "unregistered visitor" refers to a specific visitor whose facial image data is not registered in the system.

[0013] The "entrance" is the location that serves as the entrance and exit to the residence and where the camera is installed.

[0014] "Capture" refers to the action of the camera taking a picture of the visitor's face and acquiring it as image data.

[0015] "Audio output" refers to the action of transmitting the generated audio message to the outside via a speaker or the like. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[0038] System Configuration

[0039] The system mainly consists of the following components:

[0040] 1. Terminal: This terminal is installed near the entrance and includes a camera, microphone, speaker, and control device. It captures the visitor's facial image and outputs a voice message.

[0041] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. Based on the results of the judgment, it generates a voice message and sends it back to the terminal.

[0042] 3. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home or send instructions while away from home.

[0043] Program processing explanation

[0044] Below, the processing of the program for the entire system will be explained in natural language.

[0045] Image capture and transmission

[0046] The terminal captures the visitor's facial image using a camera installed at the entrance, and transmits this image data to a server via the Internet.

[0047] Facial Recognition and User Identification

[0048] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0049] Voice message output

[0050] The server generates a voice message based on the result of the judgment. Specifically, for registered users, it greets them with "Welcome back," and for unregistered visitors, it generates a warning message saying "Welcome. This area is being recorded." This voice message is then sent back to the terminal.

[0051] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[0052] Check the status from outside

[0053] Users can check the status of their homes from outside using a dedicated smartphone app, and can send requests to the server through this app.

[0054] The server collects information on the situation near the entrance in real time and sends the data back to the user, who can then check the information on their smartphone app.

[0055] Specific examples

[0056] Example 1: Registered user goes home

[0057] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[0058] Example 2: Unregistered Visitor

[0059] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[0060] Example 3: Checking the status while away from home

[0061] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[0062] The processing flow will be explained below.

[0063] Step 1:

[0064] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[0065] Step 2:

[0066] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[0067] Step 3:

[0068] The terminal transmits the captured image data to a server via the Internet.

[0069] Step 4:

[0070] The server inputs the received image data into an AI model (face recognition model).

[0071] Step 5:

[0072] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[0073] Step 6:

[0074] The server checks the identification result against an internal database to see if the corresponding user is registered.

[0075] Step 7:

[0076] The server determines the user identification result (registered user / unregistered user) and generates a voice message based on the result.

[0077] Step 8:

[0078] The server returns the generated voice message to the terminal.

[0079] Step 9:

[0080] The device outputs the received voice message from the speaker and responds appropriately to the visitor. Specifically, it greets the visitor with "Welcome back" if the visitor is a registered user, and warns the visitor that "Welcome. This area is being recorded" if the visitor is not registered.

[0081] Step 10:

[0082] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[0083] Step 11:

[0084] The server collects real-time data on the home situation (video streams and sensor information) and sends it back to the user.

[0085] Step 12:

[0086] Users can check the requested real-time information on their smartphone app.

[0087] Step 13:

[0088] Users can enter additional instructions or messages through the app and send them to the server.

[0089] Step 14:

[0090] The server transfers the received instructions and messages to the terminal.

[0091] Step 15:

[0092] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[0093] Example 1

[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0095] While crime prevention measures are now required in homes and offices, it is also important to respond appropriately to visitors. However, existing systems do not adequately distinguish between registered and unregistered visitors and automate the response. In addition, there are limited ways to check the status of the home while away from home, and there is no way to check in real time, making it difficult to balance crime prevention and hospitality.

[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0097] In this invention, the server includes means for capturing an image of the visitor using a camera installed near the entrance, means for transmitting the captured image data to a processing device, means for identifying the visitor using image analysis technology on the processing device, means for determining on the processing device whether the visitor is a registered user, and means for generating and outputting a voice message based on the determination result from the processing device. This makes it possible to automatically determine whether the visitor is a registered user or an unregistered visitor and output an appropriate voice message. Furthermore, the status of the home can be checked in real time from outside using a dedicated application, achieving both security and hospitality.

[0098] "Capture devices" are cameras or other image capture devices installed near entrances to capture images of visitors.

[0099] "Captured image data" refers to facial images and other video data of visitors captured by a camera.

[0100] A "processing device" is a device with computing resources such as a server, which analyzes image data, identifies visitors, and generates voice messages.

[0101] "Image analysis technology" is a technology that uses AI models and facial recognition software to identify visitors based on acquired image data.

[0102] A "registered user" refers to a person whose facial image data, etc. has been registered in the system in advance.

[0103] An "unregistered visitor" refers to a person whose facial image data or other data is not registered in the system.

[0104] A "voice message" is text data converted into voice data, and is the voice that the system plays to the visitor.

[0105] A "dedicated application" is software for smartphones or tablets that allows users to check the status of their home in real time while they are away from home.

[0106] "Real-time" means that data is collected and displayed close to the moment a visitor is near the entrance.

[0107] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[0108] System Configuration

[0109] The system mainly consists of the following components:

[0110] 1. Terminal: This includes a camera and audio output device installed near the entrance, as well as a control device that controls them. It is responsible for capturing facial images of visitors and outputting audio messages. Specifically, it uses a control device such as a Raspberry Pi, a Raspberry Pi camera module, and a speaker module.

[0111] 2. Server: A processing device that receives image data and identifies visitors using image analysis technology. Specifically, it performs facial recognition using TensorFlow and OpenCV libraries. It then generates a voice message based on the results of its judgment and sends it back to the device.

[0112] 3. User device: A device such as a smartphone or tablet that can be used to check the status of the home or office while away from home or to send instructions. Specifically, a dedicated application developed with Flutter is used.

[0113] Program processing explanation

[0114] The terminal uses a camera installed at the entrance to capture a facial image of the visitor. This image data is sent to a server via the Internet. The server inputs the received image data into an AI model and identifies the visitor using image analysis technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor. The server generates a voice message based on the judgment result. Specifically, if the visitor is a registered user, it greets them with "Welcome back," and if the visitor is an unregistered visitor, it generates a warning message saying "Welcome. This area is being recorded." This voice message is sent back to the terminal. The terminal then outputs the received voice message from its speaker and takes appropriate action against the visitor.

[0115] Users can check the status of their home from outside using a dedicated smartphone app. Through this app, they can send requests to the server. The server collects information about the situation near the entrance in real time and sends the data back to the user. The user can then check the information on the smartphone app.

[0116] Specific examples

[0117] Example of a registered user returning home

[0118] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[0119] Unregistered visitor example

[0120] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[0121] Example of checking the status from outside

[0122] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[0123] Example prompts for generative AI models

[0124] Example prompt sentence:

[0125] "Please generate a greeting message for the following visitor. The visitor is registered as Tanaka. ''"

[0126] Example of the resulting result:

[0127] "Welcome back, Tanaka-san. How was your day?"

[0128] Example prompt sentence:

[0129] "Generate a warning message when a visitor is not registered. ''"

[0130] Example of the resulting result:

[0131] "Welcome. We're recording here."

[0132] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0133] Step 1:

[0134] The terminal captures the visitor's facial image using a camera installed at the entrance. When a visitor enters the entrance, the sensor reacts and the camera captures the facial image. The input is the visitor's facial image, and the output is the captured image data.

[0135] Step 2:

[0136] The terminal transmits the image data captured by the imaging device to a server via the Internet. The HTTP protocol is used for transmission, and the image data is encoded and sent to the server. The input is the captured image data, and the output is the image data sent to the server.

[0137] Step 3:

[0138] The server temporarily stores the received image data and prepares it for use as input to the AI ​​model. At this time, the image data is decoded and stored in memory. The input is the image data sent to the server, and the output is image data converted into an analyzable format.

[0139] Step 4:

[0140] The server inputs image data into an AI model to perform facial recognition. The AI ​​model is trained based on facial image data of pre-registered users. Libraries such as TensorFlow and OpenCV are used for the identification process. The input is image data converted into an analyzable format, and the output is the identification result of whether the visitor is a registered user or an unregistered visitor.

[0141] Step 5:

[0142] The server generates a voice message based on the identification result. For registered users, it generates a text message such as "Welcome back," and for unregistered visitors, it converts the text message into voice data using the Google Text-to-Speech API. The input is the identification result, and the output is the generated voice data.

[0143] Step 6:

[0144] The server sends the generated audio data to the terminal, again using the HTTP protocol. The input is the generated audio data, and the output is the audio data sent to the terminal.

[0145] Step 7:

[0146] The terminal plays the received audio data to the visitor through the speaker. The terminal controls the speaker and outputs the audio message at an appropriate volume. The input is the audio data sent to the terminal, and the output is the audio message played by the speaker.

[0147] Step 8:

[0148] A user can check the status of their home from outside using a dedicated smartphone app. The user operates the app and sends requests to the server. The input is the user's request, and the output is the request sent to the server.

[0149] Step 9:

[0150] The server communicates with the terminal and collects information on the situation near the entrance in real time. The server acquires data from cameras and sensors and analyzes the real-time images and situations. The input is a request from the user, and the output is the collected real-time data.

[0151] Step 10:

[0152] The server sends the collected real-time data back to the user's smartphone app, where the user can view the data. The input is the collected real-time data, and the output is the data sent to the user's smartphone app.

[0153] (Application example 1)

[0154] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0155] In conventional factories, it has sometimes been difficult to reliably identify visitors and employees and respond appropriately. As a result, there have been many issues with factory crime prevention measures and operational efficiency. For example, there is a need to strengthen security to prevent unauthorized visitors from entering the factory without permission, and to provide prompt and appropriate hospitality to employees. The purpose of the present invention is to solve these issues and improve security and operational efficiency within the factory.

[0156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0157] In this invention, the server includes a means for identifying visitors using image analysis technology, a means for generating and outputting voice messages, and a means for controlling the voice output device according to a predetermined algorithm. This makes it possible to output a message such as "Welcome back" to registered users and a message such as "Hello. A representative will be there shortly, so please wait" to unregistered visitors. This makes it possible to strengthen crime prevention measures within the factory while providing appropriate hospitality to employees.

[0158] A "photography device" is a device installed near the entrance to capture images of visitors.

[0159] The "communication device" is a device for transmitting acquired image data to a server.

[0160] "Image analysis technology" is a technology that analyzes acquired image data and identifies visitors.

[0161] "User" is a term that refers to a pre-registered individual or employee.

[0162] "Unregistered Visitor" is a term used to refer to an individual or visitor who is not registered.

[0163] A "server" is a central processing unit for receiving and analyzing image data.

[0164] A "voice message" is a message in voice format that is generated and output based on the determination result of the server.

[0165] The "audio output device" is a device for outputting the generated audio message to the visitor.

[0166] The "predetermined algorithm" is a rule or calculation method that defines a certain procedure used by the server when controlling the audio output device.

[0167] As an embodiment of the present invention, a visitor / employee identification system for a factory will be described. This system is composed of a camera device, a communication device, a server, and an audio output device installed near the entrance.

[0168] System Configuration

[0169] 1. Imaging equipment

[0170] A camera (e.g., Logitech C920 HD Pro) installed near the entrance is responsible for capturing facial images of visitors and employees. The camera is controlled using the OpenCV library.

[0171] 2. Communications Equipment

[0172] The acquired facial image data is transmitted to a server via a communication device, using Internet Protocol to ensure secure transfer of data.

[0173] 3. Server

[0174] The server processes the image data using image analysis technology to identify visitors and employees. Specifically, it uses a pre-trained AI model (using the scikit-learn library) for facial recognition. Based on the identification results, the server generates an appropriate voice message.

[0175] 4. Audio Output Device

[0176] The voice output device outputs the generated voice messages to visitors and employees. This is done using the text_to_speech library.

[0177] Program processing explanation

[0178] Server Action:

[0179] The server first receives the image data sent from the camera. This image data is input into an AI model to identify whether the visitor is a registered employee or an unregistered visitor. Based on the identification result, a voice message such as "Welcome back" or "Hello. A representative will be there shortly, so please wait." is generated.

[0180] Audio output device handling:

[0181] The voice message received from the server is output through the voice output device. Specifically, the speak function is called to play the voice message from the speaker.

[0182] Specific examples

[0183] Example 1:

[0184] When a factory employee enters the entrance:

[0185] 1. The camera captures a facial image.

[0186] 2. The image data is sent to the server via the communication device.

[0187] 3. The server uses an AI model to identify the employee and generate a "Welcome back" message.

[0188] 4. The audio output device outputs the message.

[0189] Example 2:

[0190] When an unregistered visitor enters the entrance:

[0191] 1. The camera captures a facial image.

[0192] 2. The image data is sent to the server via the communication device.

[0193] 3. The server uses an AI model to identify the visitor as unregistered and generates a message saying, "Hello, please wait; a representative will be with you shortly."

[0194] 4. The audio output device outputs the message.

[0195] Example prompts to input to the generative AI model

[0196] "Generate a program to realize a system that identifies employees and visitors at the entrance of a factory and outputs a voice message saying "Welcome back" to employees and "Hello. A representative will be with you shortly" to visitors. Imagine an application that uses OpenCV and scikit-learn, working in conjunction with a camera and speaker."

[0197] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0198] Step 1:

[0199] The terminal (photographing device) captures the facial images of visitors who come near the entrance. The input is a real-time facial image, and the output is an image file in JPEG format or similar. This processing is performed using the OpenCV library.

[0200] Step 2:

[0201] The device sends the captured image data to a server over the Internet. The input is the captured image file, and the output is packets of transmitted data. The data is transferred securely using a communication protocol.

[0202] Step 3:

[0203] The server inputs the received image data into an AI model to identify the visitor. The input is the transferred image file, and the output is the identification result (e.g., registered employee, unregistered visitor). Scikit-learn is used as the AI ​​model, and image recognition is performed using a pre-trained model.

[0204] Step 4:

[0205] The server generates an appropriate voice message based on the identification results. The input is the identification results, and the output is a text message (e.g., "Welcome back," "Hello. A representative will be there shortly, so please wait"). If necessary, the generative AI model adjusts the details of the message.

[0206] Step 5:

[0207] The server then sends the generated voice message to the terminal's voice output device, with the input being the textual voice message and the output being packets of data to be transmitted, again using a communications protocol.

[0208] Step 6:

[0209] The terminal's audio output device converts the received voice message into speech and outputs it to the visitor through the speaker. The input is the text-based voice message, and the output is the actual voice message. This process is performed using the text_to_speech library.

[0210] Step 7:

[0211] A user (e.g., a factory manager) checks the status of their home from outside using a smartphone app. The input is a confirmation request from the user, and the output is real-time video data showing the situation near the entrance. The video is transferred from the server to the smartphone, and the user can check it using the app.

[0212] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0213] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[0214] System Configuration

[0215] The system mainly consists of the following components:

[0216] 1. Terminal: Includes a camera, microphone, speaker, and control device installed near the entrance.

[0217] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[0218] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[0219] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[0220] Program processing explanation

[0221] Below, the processing of the program for the entire system will be explained in natural language.

[0222] Image capture and transmission

[0223] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[0224] Facial Recognition and User Identification

[0225] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0226] Emotion Recognition and Voice Message Generation

[0227] The server then uses an emotion engine to analyze the visitor's emotions from the received facial images and audio data, which can detect positive, negative, or neutral emotional states.

[0228] Based on the results of the identification and emotion analysis, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, a cheerful message such as "Welcome back, you have a lovely smile today!" is generated. If a negative emotion is detected, an encouraging message such as "Welcome back, you seem a little tired, are you okay?" is generated.

[0229] The server returns the generated voice message to the terminal, and the terminal outputs the received voice message from a speaker.

[0230] Check the status from outside

[0231] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[0232] Specific examples

[0233] Example 1: Registered user goes home

[0234] A user (Mr. Tanaka) enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses facial recognition technology to identify Mr. Tanaka and uses an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message "Welcome back, you have a lovely smile today!" from its speaker.

[0235] Example 2: Unregistered Visitor

[0236] A visitor (Mr. Sato) enters the front door. The device uses a camera to capture an image of Mr. Sato's face and sends it to the server. The server uses facial recognition technology to identify Mr. Sato as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[0237] Example 3: Checking the status while away from home

[0238] While Mr. Tanaka is out, he uses an app on his smartphone to check the situation at home. He sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitor's emotion recognition results, and sends them back to Mr. Tanaka. Mr. Tanaka can check the information in the app and send additional instructions to the server from the app if necessary.

[0239] Specific steps for person recognition and emotion analysis

[0240] 1. The device captures the visitor's facial image using a camera installed at the entrance and sends it to the server.

[0241] 2. The server uses facial recognition technology to identify visitors and determine whether they are registered users or unregistered visitors.

[0242] 3. The server analyzes the identified visitor's emotions using an emotion engine.

[0243] 4. The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[0244] 5. The server returns the generated voice message to the terminal, and the terminal outputs the voice from the speaker.

[0245] In this way, the present invention is a system that not only identifies visitors but also automatically responds appropriately according to their emotions, providing both crime prevention measures and hospitality at the same time.

[0246] The processing flow will be explained below.

[0247] Step 1:

[0248] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[0249] Step 2:

[0250] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[0251] Step 3:

[0252] The terminal transmits the captured image data to a server via the Internet.

[0253] Step 4:

[0254] The server inputs the received image data into an AI model (face recognition model).

[0255] Step 5:

[0256] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[0257] Step 6:

[0258] The server checks the identification result against an internal database to see if the corresponding user is registered.

[0259] Step 7:

[0260] The server determines the user identification result (registered user / unregistered user).

[0261] Step 8:

[0262] The server inputs the facial image data of the identified visitor into the emotion engine.

[0263] Step 9:

[0264] The server uses an emotion engine to analyze the visitor's emotions and detect positive, negative, or neutral emotional states.

[0265] Step 10:

[0266] The server generates an appropriate voice message based on the user identification and emotion analysis results. For example, if a positive emotion is detected in a registered user, the server generates a message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates a message such as "Welcome back, you seem a little tired, are you okay?"

[0267] Step 11:

[0268] The server returns the generated voice message to the terminal.

[0269] Step 12:

[0270] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[0271] Step 13:

[0272] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[0273] Step 14:

[0274] The server collects real-time data on the home situation (such as video streams and emotion recognition results) and sends it back to the user.

[0275] Step 15:

[0276] Users can check the requested real-time information on their smartphone app.

[0277] Step 16:

[0278] Users can enter additional instructions or messages through the app and send them to the server.

[0279] Step 17:

[0280] The server transfers the received instructions and messages to the terminal.

[0281] Step 18:

[0282] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[0283] Example 2

[0284] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0285] Conventional visitor identification systems are limited to identifying visitors through facial recognition and have the problem of being unable to respond to visitors with consideration for their emotional state. As a result, not only are hospitality towards visitors lacking, but they are also insufficient as a crime prevention measure. The purpose of this invention is to simultaneously strengthen hospitality and crime prevention measures by identifying the emotional state of visitors and automatically responding appropriately accordingly.

[0286] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0287] In this invention, the server includes means for analyzing emotions from received facial images, means for generating and outputting an appropriate voice message based on the analyzed emotions, and means for determining whether the visitor is a registered user or an unregistered visitor, thereby enabling a response that takes into account the visitor's emotional state.

[0288] A "camera" is a photographic device installed near the entrance to capture images of visitors.

[0289] A "terminal" is a control device including a camera, and is a device that detects visitors, captures images, and transmits image data.

[0290] The "server" is an information processing device that processes received image data, identifies visitors, analyzes their emotions, and generates voice messages.

[0291] "Image data" is digital data of the visitor's facial image captured by the camera.

[0292] "Image recognition technology" is a technology for analyzing image data and identifying specific people.

[0293] A "registered user" is a visitor whose face image data has been registered in advance in the system.

[0294] A "voice message" is a voice notification message that is output to a visitor.

[0295] The "emotion engine" is a software module for analyzing emotional states from facial images and voice data.

[0296] "Emotion analysis" is the process of analyzing a visitor's facial image and voice data to determine their emotional state.

[0297] "Positive emotions" refer to positive emotional states such as joy and satisfaction.

[0298] "Negative affect" refers to negative emotional states such as sadness or dissatisfaction.

[0299] A "Secure Data Transfer Protocol" is a communications method for safely and securely transferring data over the Internet.

[0300] "Hospitality" means treating visitors with a hospitable attitude.

[0301] "Crime prevention measures" are measures to prevent crime and fraud through visitor identification and emotion analysis.

[0302] MODE FOR CARRYING OUT THE INVENTION

[0303] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[0304] System Configuration

[0305] The system mainly consists of the following components:

[0306] 1. Terminal: A device equipped with a camera, microphone, speaker, and control unit installed near the entrance. The terminal captures the visitor's facial image and transmits the data to the server.

[0307] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[0308] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[0309] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[0310] Program processing

[0311] Below, the processing of the program for the entire system will be explained in natural language.

[0312] 1. Image capture and transmission

[0313] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[0314] 2. Facial Recognition and User Identification

[0315] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0316] 3. Emotion Recognition and Voice Message Generation

[0317] Next, the server uses an emotion engine to analyze the visitor's emotions from the received facial images and voice data. The emotion engine can detect emotional states such as positive, negative, and neutral. Based on the identification and emotion analysis results, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, the server generates a cheerful message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates an encouraging message such as "Welcome back, you seem a little tired, are you okay?" The server then sends the generated voice message back to the terminal, and the terminal outputs the received voice message from its speaker.

[0318] 4. Check the status while on the go

[0319] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[0320] Specific example of system operation

[0321] Example 1: Registered user goes home

[0322] A user enters the front door. The device captures the user's facial image with a camera and sends it to the server. The server uses facial recognition technology to identify the user and an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message from its speaker: "Welcome back, you have a lovely smile today!"

[0323] Example 2: Unregistered Visitor

[0324] A visitor enters the front door. The device uses a camera to capture the visitor's facial image and sends it to the server. The server uses facial recognition technology to identify the visitor as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device then outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[0325] Example 3: Checking the status while away from home

[0326] While the user is out, they can check the status of their home using a smartphone app. The user sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitors' emotion recognition results, and sends them back to the user. The user can then check the information in the app and, if necessary, send additional instructions to the server.

[0327] Prompt Sentence Examples

[0328] A concrete example of a prompt might be:

[0329] "Facial image captured. Identification result is unregistered user. Emotion analysis results indicate the customer is in a negative emotional state."

[0330] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0331] Processing Step Description

[0332] Step 1: Visitor detection

[0333] The device detects human movement using a motion sensor installed at the entrance, and when the sensor responds, the device activates the camera.

[0334] Input: Motion sensor detection signal

[0335] Output: Camera activation signal

[0336] Specific operation: When the motion sensor is activated, the control device turns on the camera.

[0337] Step 2: Capture a face image

[0338] The device uses the activated camera to capture a facial image of the visitor.

[0339] Input: Camera status

[0340] Output: Face image data

[0341] What it does: The camera lens automatically adjusts focus and captures high-resolution images.

[0342] Step 3: Sending facial image data

[0343] The terminal transmits the captured facial image data to a server via the Internet.

[0344] Input: Face image data

[0345] Output: Image data sent to the server

[0346] What it does: The device encrypts the data and sends it securely to the server using the HTTPs protocol.

[0347] Step 4: Receiving facial image data

[0348] The server receives the face image data sent from the terminal.

[0349] Input: Image data from the device

[0350] Output: Received image data

[0351] Specific operation: Image data is temporarily stored in a database within the server for subsequent processing.

[0352] Step 5: Facial Recognition

[0353] The server inputs the received facial image data into an AI model to determine whether the visitor is a registered user or an unregistered visitor.

[0354] Input: Received image data

[0355] Output: Recognition results (registered users / unregistered visitors)

[0356] Specific operation: The server preprocesses (normalizes and resizes) the facial image, inputs it into the AI ​​model to perform facial recognition, and obtains a matching profile.

[0357] Step 6: Sentiment Analysis

[0358] The server uses an emotion engine to analyze emotions based on the image data and analysis results after facial recognition.

[0359] Input: Recognized image data

[0360] Output: Sentiment analysis result (positive / negative / neutral)

[0361] Specific operations: Extract facial feature points and analyze them to determine emotional state, and also analyze the voice spectrum if there are voice characteristics.

[0362] Step 7: Generate a voice message

[0363] The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[0364] Input: Recognition results, emotion analysis results

[0365] Output: The generated voice message

[0366] Specific operation: Using templates and analysis results, dynamically generate voice messages in text format and convert them into audio files.

[0367] Step 8: Send a voice message

[0368] The server transmits the generated voice message to the terminal.

[0369] Input: The generated voice message

[0370] Output: Audio message sent to the device

[0371] What it does: Compresses the voice message file and sends it to the device using a secure transfer protocol.

[0372] Step 9: Output a voice message

[0373] The terminal outputs the voice message received from the server from a speaker.

[0374] Input: Voice message sent to the device

[0375] Output: A voice message played through the speaker

[0376] Specific operation: The device decodes the received audio file and plays it on the internal speaker.

[0377] Step 10: Check the status on the go

[0378] Users can check the status of their home using a dedicated app.

[0379] Input: User request

[0380] Output: Real-time video and emotion recognition results

[0381] How it works: When a user sends a request from the app, the server collects the latest video data and emotion analysis results and immediately provides feedback to the user.

[0382] (Application example 2)

[0383] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0384] Conventional visitor response systems only use facial recognition technology to identify visitors, but do not take into account their emotional state. As a result, they are unable to determine the emotional state of the visitor and respond appropriately accordingly. Furthermore, there is a lack of technological means to improve hospitality for customers in brick-and-mortar stores.

[0385] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing facial image data to recognize the visitor's emotion, and means for generating and outputting an appropriate voice message based on the judgment result and emotion recognition result from the server. This makes it possible to provide individual responses according to the emotional state of the visitor, improving hospitality and providing optimal visitor service.

[0386] A "camera" is a device for capturing images and videos.

[0387] "Image data" is a digital representation of captured image information.

[0388] A "server" is a computer system that processes data and provides services over a network.

[0389] "Image recognition technology" is a technology that analyzes digital images and identifies specific objects or people.

[0390] A "visitor" refers to a person who visits a particular location.

[0391] A "registered user" is a specific person whose information is stored in advance in a database.

[0392] "Emotion recognition" is a technology that analyzes and recognizes a person's emotional state from image and audio information.

[0393] A "voice message" is a message for conveying information using voice.

[0394] The "judgment result" is the analysis result derived by the server using image recognition technology and emotion recognition technology.

[0395] The system for implementing this invention mainly captures images of visitors, identifies them based on the images, and then recognizes their emotions and outputs appropriate voice messages. This system is composed of a camera, a server, and a user terminal. Specific embodiments of the system are described below.

[0396] System Configuration

[0397] 1. Camera

[0398] The camera will be installed near the entrance and will capture facial images of visitors. The camera has high resolution and can take images in real time.

[0399] 2. Server

[0400] The server receives image data sent from the camera and identifies the visitor using image recognition technology. It then analyzes the visitor's emotions using emotion recognition technology (emotion engine). The server generates an appropriate voice message based on the identification and emotion analysis results and sends the voice message to the device.

[0401] 3. User Device

[0402] User terminals are devices such as smartphones, tablets, or smart glasses that can be used to check the system status and send instructions, including how to respond to visitors.

[0403] Program processing explanation

[0404] The server uses the following hardware and software:

[0405] Hardware

[0406] High-resolution cameras: Installed near the entrance to capture facial images of visitors.

[0407] Server: A powerful computer system for analyzing and processing data.

[0408] software

[0409] OpenCV: Used for image capture and processing.

[0410] EmotionRecognizer: A software module for analyzing emotions.

[0411] Server application: A dedicated application that receives image data, analyzes it, and generates voice messages.

[0412] The server receives image data captured by the camera and performs facial recognition using OpenCV. It then analyzes the visitor's emotions using EmotionRecognizer. Based on the analysis results, it generates an appropriate voice message and sends it back to the device to optimize visitor interaction.

[0413] Specific examples

[0414] Specific examples are given below:

[0415] Example 1: Response to a registered user's visit

[0416] 1. When a visitor approaches the entrance, a camera captures their facial image.

[0417] 2. The image data is sent to the server.

[0418] 3. The server uses facial recognition technology to identify the visitor as a registered user.

[0419] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as positive, for example.

[0420] 5. The server generates a voice message saying "Welcome back! Have a great day today!" and sends it to the device.

[0421] Example 2: How to handle unregistered visitors

[0422] 1. When a visitor approaches the entrance, a camera captures their facial image.

[0423] 2. The image data is sent to the server.

[0424] 3. The server uses facial recognition technology to identify the visitor as an unregistered person.

[0425] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as negative, for example.

[0426] 5. The server generates a voice message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it to the device.

[0427] Prompt Sentence Examples

[0428] "Taking a facial image as input, analyze the emotion as positive, negative, or neutral, and generate an appropriate message based on the analysis results."

[0429] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0430] Step 1:

[0431] A camera is installed near the entrance, and when a visitor approaches, the sensor detects their movement. The camera captures the visitor's facial image. The input here is the "visitor's facial image" and the output is the "captured visitor's facial image data."

[0432] Step 2:

[0433] The terminal sends the captured facial image data to a server via the Internet. The input here is the "captured facial image data of the visitor" and the output is the "facial image data sent to the server."

[0434] Step 3:

[0435] The server processes the received facial image data and identifies the visitor using image recognition technology. In this process, the server analyzes the facial image data using an AI model. The input is the "facial image data sent to the server" and the output is "information on the identified visitor."

[0436] Step 4:

[0437] Based on the identification result, the server refers to the database to determine whether the visitor is a registered user. At this time, it searches the database and obtains a response. The input is "identified visitor information" and the output is "the determination result of whether the user is registered or not."

[0438] Step 5:

[0439] The server uses an emotion engine based on the visitor's image data to perform emotion analysis. Emotion analysis uses a generative AI model to analyze emotions from facial expressions and other features. The input is the visitor's facial image data, and the output is the analyzed emotion data.

[0440] Step 6:

[0441] The server generates an appropriate voice message based on the results of the classification and emotion analysis. The generated voice message is based on a pre-defined prompt. The input is the "judgment result and emotion data," and the output is the "generated voice message."

[0442] Step 7:

[0443] The server sends the generated voice message to the terminal, and the terminal outputs the voice message through a speaker. Here, the input is the "generated voice message" and the output is the "voice output for the visitor."

[0444] Step 8:

[0445] A user can use a dedicated app to check the status of their home while away from home and send a request to the server. The input is a "request from the user" and the output is a "data request to the server."

[0446] Step 9:

[0447] The server collects data such as real-time video and emotion recognition results and returns them to the user. The input is an external data collection request, and the output is the data returned to the user.

[0448] Step 10:

[0449] Through the app, users can check the collected data and send additional instructions to the server if necessary. The input is "collected data" and the output is "information display and additional instructions to the user."

[0450] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0451] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0452] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0453] [Second embodiment]

[0454] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0455] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0456] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0457] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0458] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0460] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0461] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0462] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0463] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0464] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0465] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0466] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[0467] System Configuration

[0468] The system mainly consists of the following components:

[0469] 1. Terminal: This terminal is installed near the entrance and includes a camera, microphone, speaker, and control device. It captures the visitor's facial image and outputs a voice message.

[0470] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. Based on the results of the judgment, it generates a voice message and sends it back to the terminal.

[0471] 3. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home or send instructions while away from home.

[0472] Program processing explanation

[0473] Below, the processing of the program for the entire system will be explained in natural language.

[0474] Image capture and transmission

[0475] The terminal captures the visitor's facial image using a camera installed at the entrance, and transmits this image data to a server via the Internet.

[0476] Facial Recognition and User Identification

[0477] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0478] Voice message output

[0479] The server generates a voice message based on the result of the judgment. Specifically, for registered users, it greets them with "Welcome back," and for unregistered visitors, it generates a warning message saying "Welcome. This area is being recorded." This voice message is then sent back to the terminal.

[0480] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[0481] Check the status from outside

[0482] Users can check the status of their homes from outside using a dedicated smartphone app, and can send requests to the server through this app.

[0483] The server collects information on the situation near the entrance in real time and sends the data back to the user, who can then check the information on their smartphone app.

[0484] Specific examples

[0485] Example 1: Registered user goes home

[0486] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[0487] Example 2: Unregistered Visitor

[0488] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[0489] Example 3: Checking the status while away from home

[0490] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[0491] The processing flow will be explained below.

[0492] Step 1:

[0493] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[0494] Step 2:

[0495] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[0496] Step 3:

[0497] The terminal transmits the captured image data to a server via the Internet.

[0498] Step 4:

[0499] The server inputs the received image data into an AI model (face recognition model).

[0500] Step 5:

[0501] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[0502] Step 6:

[0503] The server checks the identification result against an internal database to see if the corresponding user is registered.

[0504] Step 7:

[0505] The server determines the user identification result (registered user / unregistered user) and generates a voice message based on the result.

[0506] Step 8:

[0507] The server returns the generated voice message to the terminal.

[0508] Step 9:

[0509] The device outputs the received voice message from the speaker and responds appropriately to the visitor. Specifically, it greets the visitor with "Welcome back" if the visitor is a registered user, and warns the visitor that "Welcome. This area is being recorded" if the visitor is not registered.

[0510] Step 10:

[0511] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[0512] Step 11:

[0513] The server collects real-time data on the home situation (video streams and sensor information) and sends it back to the user.

[0514] Step 12:

[0515] Users can check the requested real-time information on their smartphone app.

[0516] Step 13:

[0517] Users can enter additional instructions or messages through the app and send them to the server.

[0518] Step 14:

[0519] The server transfers the received instructions and messages to the terminal.

[0520] Step 15:

[0521] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[0522] Example 1

[0523] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0524] While crime prevention measures are now required in homes and offices, it is also important to respond appropriately to visitors. However, existing systems do not adequately distinguish between registered and unregistered visitors and automate the response. In addition, there are limited ways to check the status of the home while away from home, and there is no way to check in real time, making it difficult to balance crime prevention and hospitality.

[0525] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0526] In this invention, the server includes means for capturing an image of the visitor using a camera installed near the entrance, means for transmitting the captured image data to a processing device, means for identifying the visitor using image analysis technology on the processing device, means for determining on the processing device whether the visitor is a registered user, and means for generating and outputting a voice message based on the determination result from the processing device. This makes it possible to automatically determine whether the visitor is a registered user or an unregistered visitor and output an appropriate voice message. Furthermore, the status of the home can be checked in real time from outside using a dedicated application, achieving both security and hospitality.

[0527] "Capture devices" are cameras or other image capture devices installed near entrances to capture images of visitors.

[0528] "Captured image data" refers to facial images and other video data of visitors captured by a camera.

[0529] A "processing device" is a device with computing resources such as a server, which analyzes image data, identifies visitors, and generates voice messages.

[0530] "Image analysis technology" is a technology that uses AI models and facial recognition software to identify visitors based on acquired image data.

[0531] A "registered user" refers to a person whose facial image data, etc. has been registered in the system in advance.

[0532] An "unregistered visitor" refers to a person whose facial image data or other data is not registered in the system.

[0533] A "voice message" is text data converted into voice data, and is the voice that the system plays to the visitor.

[0534] A "dedicated application" is software for smartphones or tablets that allows users to check the status of their home in real time while they are away from home.

[0535] "Real-time" means that data is collected and displayed close to the moment a visitor is near the entrance.

[0536] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[0537] System Configuration

[0538] The system mainly consists of the following components:

[0539] 1. Terminal: This includes a camera and audio output device installed near the entrance, as well as a control device that controls them. It is responsible for capturing facial images of visitors and outputting audio messages. Specifically, it uses a control device such as a Raspberry Pi, a Raspberry Pi camera module, and a speaker module.

[0540] 2. Server: A processing device that receives image data and identifies visitors using image analysis technology. Specifically, it performs facial recognition using TensorFlow and OpenCV libraries. It then generates a voice message based on the results of its judgment and sends it back to the device.

[0541] 3. User device: A device such as a smartphone or tablet that can be used to check the status of the home or office while away from home or to send instructions. Specifically, a dedicated application developed with Flutter is used.

[0542] Program processing explanation

[0543] The terminal uses a camera installed at the entrance to capture a facial image of the visitor. This image data is sent to a server via the Internet. The server inputs the received image data into an AI model and identifies the visitor using image analysis technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor. The server generates a voice message based on the judgment result. Specifically, if the visitor is a registered user, it greets them with "Welcome back," and if the visitor is an unregistered visitor, it generates a warning message saying "Welcome. This area is being recorded." This voice message is sent back to the terminal. The terminal then outputs the received voice message from its speaker and takes appropriate action against the visitor.

[0544] Users can check the status of their home from outside using a dedicated smartphone app. Through this app, they can send requests to the server. The server collects information about the situation near the entrance in real time and sends the data back to the user. The user can then check the information on the smartphone app.

[0545] Specific examples

[0546] Example of a registered user returning home

[0547] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[0548] Unregistered visitor example

[0549] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[0550] Example of checking the status from outside

[0551] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[0552] Example prompts for generative AI models

[0553] Example prompt sentence:

[0554] "Please generate a greeting message for the following visitor. The visitor is registered as Tanaka. ''"

[0555] Example of the resulting result:

[0556] "Welcome back, Tanaka-san. How was your day?"

[0557] Example prompt sentence:

[0558] "Generate a warning message when a visitor is not registered. ''"

[0559] Example of the resulting result:

[0560] "Welcome. We're recording here."

[0561] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0562] Step 1:

[0563] The terminal captures the visitor's facial image using a camera installed at the entrance. When a visitor enters the entrance, the sensor reacts and the camera captures the facial image. The input is the visitor's facial image, and the output is the captured image data.

[0564] Step 2:

[0565] The terminal transmits the image data captured by the imaging device to a server via the Internet. The HTTP protocol is used for transmission, and the image data is encoded and sent to the server. The input is the captured image data, and the output is the image data sent to the server.

[0566] Step 3:

[0567] The server temporarily stores the received image data and prepares it for use as input to the AI ​​model. At this time, the image data is decoded and stored in memory. The input is the image data sent to the server, and the output is image data converted into an analyzable format.

[0568] Step 4:

[0569] The server inputs image data into an AI model to perform facial recognition. The AI ​​model is trained based on facial image data of pre-registered users. Libraries such as TensorFlow and OpenCV are used for the identification process. The input is image data converted into an analyzable format, and the output is the identification result of whether the visitor is a registered user or an unregistered visitor.

[0570] Step 5:

[0571] The server generates a voice message based on the identification result. For registered users, it generates a text message such as "Welcome back," and for unregistered visitors, it converts the text message into voice data using the Google Text-to-Speech API. The input is the identification result, and the output is the generated voice data.

[0572] Step 6:

[0573] The server sends the generated audio data to the terminal, again using the HTTP protocol. The input is the generated audio data, and the output is the audio data sent to the terminal.

[0574] Step 7:

[0575] The terminal plays the received audio data to the visitor through the speaker. The terminal controls the speaker and outputs the audio message at an appropriate volume. The input is the audio data sent to the terminal, and the output is the audio message played by the speaker.

[0576] Step 8:

[0577] A user can check the status of their home from outside using a dedicated smartphone app. The user operates the app and sends requests to the server. The input is the user's request, and the output is the request sent to the server.

[0578] Step 9:

[0579] The server communicates with the terminal and collects information on the situation near the entrance in real time. The server acquires data from cameras and sensors and analyzes the real-time images and situations. The input is a request from the user, and the output is the collected real-time data.

[0580] Step 10:

[0581] The server sends the collected real-time data back to the user's smartphone app, where the user can view the data. The input is the collected real-time data, and the output is the data sent to the user's smartphone app.

[0582] (Application example 1)

[0583] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0584] In conventional factories, it has sometimes been difficult to reliably identify visitors and employees and respond appropriately. As a result, there have been many issues with factory crime prevention measures and operational efficiency. For example, there is a need to strengthen security to prevent unauthorized visitors from entering the factory without permission, and to provide prompt and appropriate hospitality to employees. The purpose of the present invention is to solve these issues and improve security and operational efficiency within the factory.

[0585] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0586] In this invention, the server includes a means for identifying visitors using image analysis technology, a means for generating and outputting voice messages, and a means for controlling the voice output device according to a predetermined algorithm. This makes it possible to output a message such as "Welcome back" to registered users and a message such as "Hello. A representative will be there shortly, so please wait" to unregistered visitors. This makes it possible to strengthen crime prevention measures within the factory while providing appropriate hospitality to employees.

[0587] A "photography device" is a device installed near the entrance to capture images of visitors.

[0588] The "communication device" is a device for transmitting acquired image data to a server.

[0589] "Image analysis technology" is a technology that analyzes acquired image data and identifies visitors.

[0590] "User" is a term that refers to a pre-registered individual or employee.

[0591] "Unregistered Visitor" is a term used to refer to an individual or visitor who is not registered.

[0592] A "server" is a central processing unit for receiving and analyzing image data.

[0593] A "voice message" is a message in voice format that is generated and output based on the determination result of the server.

[0594] The "audio output device" is a device for outputting the generated audio message to the visitor.

[0595] The "predetermined algorithm" is a rule or calculation method that defines a certain procedure used by the server when controlling the audio output device.

[0596] As an embodiment of the present invention, a visitor / employee identification system for a factory will be described. This system is composed of a camera device, a communication device, a server, and an audio output device installed near the entrance.

[0597] System Configuration

[0598] 1. Imaging equipment

[0599] A camera (e.g., Logitech C920 HD Pro) installed near the entrance is responsible for capturing facial images of visitors and employees. The camera is controlled using the OpenCV library.

[0600] 2. Communications Equipment

[0601] The acquired facial image data is transmitted to a server via a communication device, using Internet Protocol to ensure secure transfer of data.

[0602] 3. Server

[0603] The server processes the image data using image analysis technology to identify visitors and employees. Specifically, it uses a pre-trained AI model (using the scikit-learn library) for facial recognition. Based on the identification results, the server generates an appropriate voice message.

[0604] 4. Audio Output Device

[0605] The voice output device outputs the generated voice messages to visitors and employees. This is done using the text_to_speech library.

[0606] Program processing explanation

[0607] Server Action:

[0608] The server first receives the image data sent from the camera. This image data is input into an AI model to identify whether the visitor is a registered employee or an unregistered visitor. Based on the identification result, a voice message such as "Welcome back" or "Hello. A representative will be there shortly, so please wait." is generated.

[0609] Audio output device handling:

[0610] The voice message received from the server is output through the voice output device. Specifically, the speak function is called to play the voice message from the speaker.

[0611] Specific examples

[0612] Example 1:

[0613] When a factory employee enters the entrance:

[0614] 1. The camera captures a facial image.

[0615] 2. The image data is sent to the server via the communication device.

[0616] 3. The server uses an AI model to identify the employee and generate a "Welcome back" message.

[0617] 4. The audio output device outputs the message.

[0618] Example 2:

[0619] When an unregistered visitor enters the entrance:

[0620] 1. The camera captures a facial image.

[0621] 2. The image data is sent to the server via the communication device.

[0622] 3. The server uses an AI model to identify the visitor as unregistered and generates a message saying, "Hello, please wait; a representative will be with you shortly."

[0623] 4. The audio output device outputs the message.

[0624] Example prompts to input to the generative AI model

[0625] "Generate a program to realize a system that identifies employees and visitors at the entrance of a factory and outputs a voice message saying "Welcome back" to employees and "Hello. A representative will be with you shortly" to visitors. Imagine an application that uses OpenCV and scikit-learn, working in conjunction with a camera and speaker."

[0626] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0627] Step 1:

[0628] The terminal (photographing device) captures the facial images of visitors who come near the entrance. The input is a real-time facial image, and the output is an image file in JPEG format or similar. This processing is performed using the OpenCV library.

[0629] Step 2:

[0630] The device sends the captured image data to a server over the Internet. The input is the captured image file, and the output is packets of transmitted data. The data is transferred securely using a communication protocol.

[0631] Step 3:

[0632] The server inputs the received image data into an AI model to identify the visitor. The input is the transferred image file, and the output is the identification result (e.g., registered employee, unregistered visitor). Scikit-learn is used as the AI ​​model, and image recognition is performed using a pre-trained model.

[0633] Step 4:

[0634] The server generates an appropriate voice message based on the identification results. The input is the identification results, and the output is a text message (e.g., "Welcome back," "Hello. A representative will be there shortly, so please wait"). If necessary, the generative AI model adjusts the details of the message.

[0635] Step 5:

[0636] The server then sends the generated voice message to the terminal's voice output device, with the input being the textual voice message and the output being packets of data to be transmitted, again using a communications protocol.

[0637] Step 6:

[0638] The terminal's audio output device converts the received voice message into speech and outputs it to the visitor through the speaker. The input is the text-based voice message, and the output is the actual voice message. This process is performed using the text_to_speech library.

[0639] Step 7:

[0640] A user (e.g., a factory manager) checks the status of their home from outside using a smartphone app. The input is a confirmation request from the user, and the output is real-time video data showing the situation near the entrance. The video is transferred from the server to the smartphone, and the user can check it using the app.

[0641] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0642] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[0643] System Configuration

[0644] The system mainly consists of the following components:

[0645] 1. Terminal: Includes a camera, microphone, speaker, and control device installed near the entrance.

[0646] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[0647] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[0648] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[0649] Program processing explanation

[0650] Below, the processing of the program for the entire system will be explained in natural language.

[0651] Image capture and transmission

[0652] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[0653] Facial Recognition and User Identification

[0654] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0655] Emotion Recognition and Voice Message Generation

[0656] The server then uses an emotion engine to analyze the visitor's emotions from the received facial images and audio data, which can detect positive, negative, or neutral emotional states.

[0657] Based on the results of the identification and emotion analysis, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, a cheerful message such as "Welcome back, you have a lovely smile today!" is generated. If a negative emotion is detected, an encouraging message such as "Welcome back, you seem a little tired, are you okay?" is generated.

[0658] The server returns the generated voice message to the terminal, and the terminal outputs the received voice message from a speaker.

[0659] Check the status from outside

[0660] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[0661] Specific examples

[0662] Example 1: Registered user goes home

[0663] A user (Mr. Tanaka) enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses facial recognition technology to identify Mr. Tanaka and uses an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message "Welcome back, you have a lovely smile today!" from its speaker.

[0664] Example 2: Unregistered Visitor

[0665] A visitor (Mr. Sato) enters the front door. The device uses a camera to capture an image of Mr. Sato's face and sends it to the server. The server uses facial recognition technology to identify Mr. Sato as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[0666] Example 3: Checking the status while away from home

[0667] While Mr. Tanaka is out, he uses an app on his smartphone to check the situation at home. He sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitor's emotion recognition results, and sends them back to Mr. Tanaka. Mr. Tanaka can check the information in the app and send additional instructions to the server from the app if necessary.

[0668] Specific steps for person recognition and emotion analysis

[0669] 1. The device captures the visitor's facial image using a camera installed at the entrance and sends it to the server.

[0670] 2. The server uses facial recognition technology to identify visitors and determine whether they are registered users or unregistered visitors.

[0671] 3. The server analyzes the identified visitor's emotions using an emotion engine.

[0672] 4. The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[0673] 5. The server returns the generated voice message to the terminal, and the terminal outputs the voice from the speaker.

[0674] In this way, the present invention is a system that not only identifies visitors but also automatically responds appropriately according to their emotions, providing both crime prevention measures and hospitality at the same time.

[0675] The processing flow will be explained below.

[0676] Step 1:

[0677] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[0678] Step 2:

[0679] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[0680] Step 3:

[0681] The terminal transmits the captured image data to a server via the Internet.

[0682] Step 4:

[0683] The server inputs the received image data into an AI model (face recognition model).

[0684] Step 5:

[0685] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[0686] Step 6:

[0687] The server checks the identification result against an internal database to see if the corresponding user is registered.

[0688] Step 7:

[0689] The server determines the user identification result (registered user / unregistered user).

[0690] Step 8:

[0691] The server inputs the facial image data of the identified visitor into the emotion engine.

[0692] Step 9:

[0693] The server uses an emotion engine to analyze the visitor's emotions and detect positive, negative, or neutral emotional states.

[0694] Step 10:

[0695] The server generates an appropriate voice message based on the user identification and emotion analysis results. For example, if a positive emotion is detected in a registered user, the server generates a message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates a message such as "Welcome back, you seem a little tired, are you okay?"

[0696] Step 11:

[0697] The server returns the generated voice message to the terminal.

[0698] Step 12:

[0699] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[0700] Step 13:

[0701] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[0702] Step 14:

[0703] The server collects real-time data on the home situation (such as video streams and emotion recognition results) and sends it back to the user.

[0704] Step 15:

[0705] Users can check the requested real-time information on their smartphone app.

[0706] Step 16:

[0707] Users can enter additional instructions or messages through the app and send them to the server.

[0708] Step 17:

[0709] The server transfers the received instructions and messages to the terminal.

[0710] Step 18:

[0711] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[0712] Example 2

[0713] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0714] Conventional visitor identification systems are limited to identifying visitors through facial recognition and have the problem of being unable to respond to visitors with consideration for their emotional state. As a result, not only are hospitality towards visitors lacking, but they are also insufficient as a crime prevention measure. The purpose of this invention is to simultaneously strengthen hospitality and crime prevention measures by identifying the emotional state of visitors and automatically responding appropriately accordingly.

[0715] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0716] In this invention, the server includes means for analyzing emotions from received facial images, means for generating and outputting an appropriate voice message based on the analyzed emotions, and means for determining whether the visitor is a registered user or an unregistered visitor, thereby enabling a response that takes into account the visitor's emotional state.

[0717] A "camera" is a photographic device installed near the entrance to capture images of visitors.

[0718] A "terminal" is a control device including a camera, and is a device that detects visitors, captures images, and transmits image data.

[0719] The "server" is an information processing device that processes received image data, identifies visitors, analyzes their emotions, and generates voice messages.

[0720] "Image data" is digital data of the visitor's facial image captured by the camera.

[0721] "Image recognition technology" is a technology for analyzing image data and identifying specific people.

[0722] A "registered user" is a visitor whose face image data has been registered in advance in the system.

[0723] A "voice message" is a voice notification message that is output to a visitor.

[0724] The "emotion engine" is a software module for analyzing emotional states from facial images and voice data.

[0725] "Emotion analysis" is the process of analyzing a visitor's facial image and voice data to determine their emotional state.

[0726] "Positive emotions" refer to positive emotional states such as joy and satisfaction.

[0727] "Negative affect" refers to negative emotional states such as sadness or dissatisfaction.

[0728] A "Secure Data Transfer Protocol" is a communications method for safely and securely transferring data over the Internet.

[0729] "Hospitality" means treating visitors with a hospitable attitude.

[0730] "Crime prevention measures" are measures to prevent crime and fraud through visitor identification and emotion analysis.

[0731] MODE FOR CARRYING OUT THE INVENTION

[0732] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[0733] System Configuration

[0734] The system mainly consists of the following components:

[0735] 1. Terminal: A device equipped with a camera, microphone, speaker, and control unit installed near the entrance. The terminal captures the visitor's facial image and transmits the data to the server.

[0736] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[0737] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[0738] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[0739] Program processing

[0740] Below, the processing of the program for the entire system will be explained in natural language.

[0741] 1. Image capture and transmission

[0742] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[0743] 2. Facial Recognition and User Identification

[0744] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0745] 3. Emotion Recognition and Voice Message Generation

[0746] Next, the server uses an emotion engine to analyze the visitor's emotions from the received facial images and voice data. The emotion engine can detect emotional states such as positive, negative, and neutral. Based on the identification and emotion analysis results, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, the server generates a cheerful message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates an encouraging message such as "Welcome back, you seem a little tired, are you okay?" The server then sends the generated voice message back to the terminal, and the terminal outputs the received voice message from its speaker.

[0747] 4. Check the status while on the go

[0748] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[0749] Specific example of system operation

[0750] Example 1: Registered user goes home

[0751] A user enters the front door. The device captures the user's facial image with a camera and sends it to the server. The server uses facial recognition technology to identify the user and an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message from its speaker: "Welcome back, you have a lovely smile today!"

[0752] Example 2: Unregistered Visitor

[0753] A visitor enters the front door. The device uses a camera to capture the visitor's facial image and sends it to the server. The server uses facial recognition technology to identify the visitor as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device then outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[0754] Example 3: Checking the status while away from home

[0755] While the user is out, they can check the status of their home using a smartphone app. The user sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitors' emotion recognition results, and sends them back to the user. The user can then check the information in the app and, if necessary, send additional instructions to the server.

[0756] Prompt Sentence Examples

[0757] A concrete example of a prompt might be:

[0758] "Facial image captured. Identification result is unregistered user. Emotion analysis results indicate the customer is in a negative emotional state."

[0759] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0760] Processing Step Description

[0761] Step 1: Visitor detection

[0762] The device detects human movement using a motion sensor installed at the entrance, and when the sensor responds, the device activates the camera.

[0763] Input: Motion sensor detection signal

[0764] Output: Camera activation signal

[0765] Specific operation: When the motion sensor is activated, the control device turns on the camera.

[0766] Step 2: Capture a face image

[0767] The device uses the activated camera to capture a facial image of the visitor.

[0768] Input: Camera status

[0769] Output: Face image data

[0770] What it does: The camera lens automatically adjusts focus and captures high-resolution images.

[0771] Step 3: Sending facial image data

[0772] The terminal transmits the captured facial image data to a server via the Internet.

[0773] Input: Face image data

[0774] Output: Image data sent to the server

[0775] What it does: The device encrypts the data and sends it securely to the server using the HTTPs protocol.

[0776] Step 4: Receiving facial image data

[0777] The server receives the face image data sent from the terminal.

[0778] Input: Image data from the device

[0779] Output: Received image data

[0780] Specific operation: Image data is temporarily stored in a database within the server for subsequent processing.

[0781] Step 5: Facial Recognition

[0782] The server inputs the received facial image data into an AI model to determine whether the visitor is a registered user or an unregistered visitor.

[0783] Input: Received image data

[0784] Output: Recognition results (registered users / unregistered visitors)

[0785] Specific operation: The server preprocesses (normalizes and resizes) the facial image, inputs it into the AI ​​model to perform facial recognition, and obtains a matching profile.

[0786] Step 6: Sentiment Analysis

[0787] The server uses an emotion engine to analyze emotions based on the image data and analysis results after facial recognition.

[0788] Input: Recognized image data

[0789] Output: Sentiment analysis result (positive / negative / neutral)

[0790] Specific operations: Extract facial feature points and analyze them to determine emotional state, and also analyze the voice spectrum if there are voice characteristics.

[0791] Step 7: Generate a voice message

[0792] The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[0793] Input: Recognition results, emotion analysis results

[0794] Output: The generated voice message

[0795] Specific operation: Using templates and analysis results, dynamically generate voice messages in text format and convert them into audio files.

[0796] Step 8: Send a voice message

[0797] The server transmits the generated voice message to the terminal.

[0798] Input: The generated voice message

[0799] Output: Audio message sent to the device

[0800] What it does: Compresses the voice message file and sends it to the device using a secure transfer protocol.

[0801] Step 9: Output a voice message

[0802] The terminal outputs the voice message received from the server from a speaker.

[0803] Input: Voice message sent to the device

[0804] Output: A voice message played through the speaker

[0805] Specific operation: The device decodes the received audio file and plays it on the internal speaker.

[0806] Step 10: Check the status on the go

[0807] Users can check the status of their home using a dedicated app.

[0808] Input: User request

[0809] Output: Real-time video and emotion recognition results

[0810] How it works: When a user sends a request from the app, the server collects the latest video data and emotion analysis results and immediately provides feedback to the user.

[0811] (Application example 2)

[0812] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0813] Conventional visitor response systems only use facial recognition technology to identify visitors, but do not take into account their emotional state. As a result, they are unable to determine the emotional state of the visitor and respond appropriately accordingly. Furthermore, there is a lack of technological means to improve hospitality for customers in brick-and-mortar stores.

[0814] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing facial image data to recognize the visitor's emotion, and means for generating and outputting an appropriate voice message based on the judgment result and emotion recognition result from the server. This makes it possible to provide individual responses according to the emotional state of the visitor, improving hospitality and providing optimal visitor service.

[0815] A "camera" is a device for capturing images and videos.

[0816] "Image data" is a digital representation of captured image information.

[0817] A "server" is a computer system that processes data and provides services over a network.

[0818] "Image recognition technology" is a technology that analyzes digital images and identifies specific objects or people.

[0819] A "visitor" refers to a person who visits a particular location.

[0820] A "registered user" is a specific person whose information is stored in advance in a database.

[0821] "Emotion recognition" is a technology that analyzes and recognizes a person's emotional state from image and audio information.

[0822] A "voice message" is a message for conveying information using voice.

[0823] The "judgment result" is the analysis result derived by the server using image recognition technology and emotion recognition technology.

[0824] The system for implementing this invention mainly captures images of visitors, identifies them based on the images, and then recognizes their emotions and outputs appropriate voice messages. This system is composed of a camera, a server, and a user terminal. Specific embodiments of the system are described below.

[0825] System Configuration

[0826] 1. Camera

[0827] The camera will be installed near the entrance and will capture facial images of visitors. The camera has high resolution and can take images in real time.

[0828] 2. Server

[0829] The server receives image data sent from the camera and identifies the visitor using image recognition technology. It then analyzes the visitor's emotions using emotion recognition technology (emotion engine). The server generates an appropriate voice message based on the identification and emotion analysis results and sends the voice message to the device.

[0830] 3. User Device

[0831] User terminals are devices such as smartphones, tablets, or smart glasses that can be used to check the system status and send instructions, including how to respond to visitors.

[0832] Program processing explanation

[0833] The server uses the following hardware and software:

[0834] Hardware

[0835] High-resolution cameras: Installed near the entrance to capture facial images of visitors.

[0836] Server: A powerful computer system for analyzing and processing data.

[0837] software

[0838] OpenCV: Used for image capture and processing.

[0839] EmotionRecognizer: A software module for analyzing emotions.

[0840] Server application: A dedicated application that receives image data, analyzes it, and generates voice messages.

[0841] The server receives image data captured by the camera and performs facial recognition using OpenCV. It then analyzes the visitor's emotions using EmotionRecognizer. Based on the analysis results, it generates an appropriate voice message and sends it back to the device to optimize visitor interaction.

[0842] Specific examples

[0843] Specific examples are given below:

[0844] Example 1: Response to a registered user's visit

[0845] 1. When a visitor approaches the entrance, a camera captures their facial image.

[0846] 2. The image data is sent to the server.

[0847] 3. The server uses facial recognition technology to identify the visitor as a registered user.

[0848] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as positive, for example.

[0849] 5. The server generates a voice message saying "Welcome back! Have a great day today!" and sends it to the device.

[0850] Example 2: How to handle unregistered visitors

[0851] 1. When a visitor approaches the entrance, a camera captures their facial image.

[0852] 2. The image data is sent to the server.

[0853] 3. The server uses facial recognition technology to identify the visitor as an unregistered person.

[0854] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as negative, for example.

[0855] 5. The server generates a voice message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it to the device.

[0856] Prompt Sentence Examples

[0857] "Taking a facial image as input, analyze the emotion as positive, negative, or neutral, and generate an appropriate message based on the analysis results."

[0858] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0859] Step 1:

[0860] A camera is installed near the entrance, and when a visitor approaches, the sensor detects their movement. The camera captures the visitor's facial image. The input here is the "visitor's facial image" and the output is the "captured visitor's facial image data."

[0861] Step 2:

[0862] The terminal sends the captured facial image data to a server via the Internet. The input here is the "captured facial image data of the visitor" and the output is the "facial image data sent to the server."

[0863] Step 3:

[0864] The server processes the received facial image data and identifies the visitor using image recognition technology. In this process, the server analyzes the facial image data using an AI model. The input is the "facial image data sent to the server" and the output is "information on the identified visitor."

[0865] Step 4:

[0866] Based on the identification result, the server refers to the database to determine whether the visitor is a registered user. At this time, it searches the database and obtains a response. The input is "identified visitor information" and the output is "the determination result of whether the user is registered or not."

[0867] Step 5:

[0868] The server uses an emotion engine based on the visitor's image data to perform emotion analysis. Emotion analysis uses a generative AI model to analyze emotions from facial expressions and other features. The input is the visitor's facial image data, and the output is the analyzed emotion data.

[0869] Step 6:

[0870] The server generates an appropriate voice message based on the results of the classification and emotion analysis. The generated voice message is based on a pre-defined prompt. The input is the "judgment result and emotion data," and the output is the "generated voice message."

[0871] Step 7:

[0872] The server sends the generated voice message to the terminal, and the terminal outputs the voice message through a speaker. Here, the input is the "generated voice message" and the output is the "voice output for the visitor."

[0873] Step 8:

[0874] A user can use a dedicated app to check the status of their home while away from home and send a request to the server. The input is a "request from the user" and the output is a "data request to the server."

[0875] Step 9:

[0876] The server collects data such as real-time video and emotion recognition results and returns them to the user. The input is an external data collection request, and the output is the data returned to the user.

[0877] Step 10:

[0878] Through the app, users can check the collected data and send additional instructions to the server if necessary. The input is "collected data" and the output is "information display and additional instructions to the user."

[0879] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0880] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0881] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0882] [Third embodiment]

[0883] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0884] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0885] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0886] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0887] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0888] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0889] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0890] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0891] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0892] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0893] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0894] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0895] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[0896] System Configuration

[0897] The system mainly consists of the following components:

[0898] 1. Terminal: This terminal is installed near the entrance and includes a camera, microphone, speaker, and control device. It captures the visitor's facial image and outputs a voice message.

[0899] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. Based on the results of the judgment, it generates a voice message and sends it back to the terminal.

[0900] 3. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home or send instructions while away from home.

[0901] Program processing explanation

[0902] Below, the processing of the program for the entire system will be explained in natural language.

[0903] Image capture and transmission

[0904] The terminal captures the visitor's facial image using a camera installed at the entrance, and transmits this image data to a server via the Internet.

[0905] Facial Recognition and User Identification

[0906] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[0907] Voice message output

[0908] The server generates a voice message based on the result of the judgment. Specifically, for registered users, it greets them with "Welcome back," and for unregistered visitors, it generates a warning message saying "Welcome. This area is being recorded." This voice message is then sent back to the terminal.

[0909] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[0910] Check the status from outside

[0911] Users can check the status of their homes from outside using a dedicated smartphone app, and can send requests to the server through this app.

[0912] The server collects information on the situation near the entrance in real time and sends the data back to the user, who can then check the information on their smartphone app.

[0913] Specific examples

[0914] Example 1: Registered user goes home

[0915] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[0916] Example 2: Unregistered Visitor

[0917] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[0918] Example 3: Checking the status while away from home

[0919] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[0920] The processing flow will be explained below.

[0921] Step 1:

[0922] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[0923] Step 2:

[0924] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[0925] Step 3:

[0926] The terminal transmits the captured image data to a server via the Internet.

[0927] Step 4:

[0928] The server inputs the received image data into an AI model (face recognition model).

[0929] Step 5:

[0930] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[0931] Step 6:

[0932] The server checks the identification result against an internal database to see if the corresponding user is registered.

[0933] Step 7:

[0934] The server determines the user identification result (registered user / unregistered user) and generates a voice message based on the result.

[0935] Step 8:

[0936] The server returns the generated voice message to the terminal.

[0937] Step 9:

[0938] The device outputs the received voice message from the speaker and responds appropriately to the visitor. Specifically, it greets the visitor with "Welcome back" if the visitor is a registered user, and warns the visitor that "Welcome. This area is being recorded" if the visitor is not registered.

[0939] Step 10:

[0940] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[0941] Step 11:

[0942] The server collects real-time data on the home situation (video streams and sensor information) and sends it back to the user.

[0943] Step 12:

[0944] Users can check the requested real-time information on their smartphone app.

[0945] Step 13:

[0946] Users can enter additional instructions or messages through the app and send them to the server.

[0947] Step 14:

[0948] The server transfers the received instructions and messages to the terminal.

[0949] Step 15:

[0950] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[0951] Example 1

[0952] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0953] While crime prevention measures are now required in homes and offices, it is also important to respond appropriately to visitors. However, existing systems do not adequately distinguish between registered and unregistered visitors and automate the response. In addition, there are limited ways to check the status of the home while away from home, and there is no way to check in real time, making it difficult to balance crime prevention and hospitality.

[0954] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0955] In this invention, the server includes means for capturing an image of the visitor using a camera installed near the entrance, means for transmitting the captured image data to a processing device, means for identifying the visitor using image analysis technology on the processing device, means for determining on the processing device whether the visitor is a registered user, and means for generating and outputting a voice message based on the determination result from the processing device. This makes it possible to automatically determine whether the visitor is a registered user or an unregistered visitor and output an appropriate voice message. Furthermore, the status of the home can be checked in real time from outside using a dedicated application, achieving both security and hospitality.

[0956] "Capture devices" are cameras or other image capture devices installed near entrances to capture images of visitors.

[0957] "Captured image data" refers to facial images and other video data of visitors captured by a camera.

[0958] A "processing device" is a device with computing resources such as a server, which analyzes image data, identifies visitors, and generates voice messages.

[0959] "Image analysis technology" is a technology that uses AI models and facial recognition software to identify visitors based on acquired image data.

[0960] A "registered user" refers to a person whose facial image data, etc. has been registered in the system in advance.

[0961] An "unregistered visitor" refers to a person whose facial image data or other data is not registered in the system.

[0962] A "voice message" is text data converted into voice data, and is the voice that the system plays to the visitor.

[0963] A "dedicated application" is software for smartphones or tablets that allows users to check the status of their home in real time while they are away from home.

[0964] "Real-time" means that data is collected and displayed close to the moment a visitor is near the entrance.

[0965] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[0966] System Configuration

[0967] The system mainly consists of the following components:

[0968] 1. Terminal: This includes a camera and audio output device installed near the entrance, as well as a control device that controls them. It is responsible for capturing facial images of visitors and outputting audio messages. Specifically, it uses a control device such as a Raspberry Pi, a Raspberry Pi camera module, and a speaker module.

[0969] 2. Server: A processing device that receives image data and identifies visitors using image analysis technology. Specifically, it performs facial recognition using TensorFlow and OpenCV libraries. It then generates a voice message based on the results of its judgment and sends it back to the device.

[0970] 3. User device: A device such as a smartphone or tablet that can be used to check the status of the home or office while away from home or to send instructions. Specifically, a dedicated application developed with Flutter is used.

[0971] Program processing explanation

[0972] The terminal uses a camera installed at the entrance to capture a facial image of the visitor. This image data is sent to a server via the Internet. The server inputs the received image data into an AI model and identifies the visitor using image analysis technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor. The server generates a voice message based on the judgment result. Specifically, if the visitor is a registered user, it greets them with "Welcome back," and if the visitor is an unregistered visitor, it generates a warning message saying "Welcome. This area is being recorded." This voice message is sent back to the terminal. The terminal then outputs the received voice message from its speaker and takes appropriate action against the visitor.

[0973] Users can check the status of their home from outside using a dedicated smartphone app. Through this app, they can send requests to the server. The server collects information about the situation near the entrance in real time and sends the data back to the user. The user can then check the information on the smartphone app.

[0974] Specific examples

[0975] Example of a registered user returning home

[0976] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[0977] Unregistered visitor example

[0978] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[0979] Example of checking the status from outside

[0980] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[0981] Example prompts for generative AI models

[0982] Example prompt sentence:

[0983] "Please generate a greeting message for the following visitor. The visitor is registered as Tanaka. ''"

[0984] Example of the resulting result:

[0985] "Welcome back, Tanaka-san. How was your day?"

[0986] Example prompt sentence:

[0987] "Generate a warning message when a visitor is not registered. ''"

[0988] Example of the resulting result:

[0989] "Welcome. We're recording here."

[0990] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0991] Step 1:

[0992] The terminal captures the visitor's facial image using a camera installed at the entrance. When a visitor enters the entrance, the sensor reacts and the camera captures the facial image. The input is the visitor's facial image, and the output is the captured image data.

[0993] Step 2:

[0994] The terminal transmits the image data captured by the imaging device to a server via the Internet. The HTTP protocol is used for transmission, and the image data is encoded and sent to the server. The input is the captured image data, and the output is the image data sent to the server.

[0995] Step 3:

[0996] The server temporarily stores the received image data and prepares it for use as input to the AI ​​model. At this time, the image data is decoded and stored in memory. The input is the image data sent to the server, and the output is image data converted into an analyzable format.

[0997] Step 4:

[0998] The server inputs image data into an AI model to perform facial recognition. The AI ​​model is trained based on facial image data of pre-registered users. Libraries such as TensorFlow and OpenCV are used for the identification process. The input is image data converted into an analyzable format, and the output is the identification result of whether the visitor is a registered user or an unregistered visitor.

[0999] Step 5:

[1000] The server generates a voice message based on the identification result. For registered users, it generates a text message such as "Welcome back," and for unregistered visitors, it converts the text message into voice data using the Google Text-to-Speech API. The input is the identification result, and the output is the generated voice data.

[1001] Step 6:

[1002] The server sends the generated audio data to the terminal, again using the HTTP protocol. The input is the generated audio data, and the output is the audio data sent to the terminal.

[1003] Step 7:

[1004] The terminal plays the received audio data to the visitor through the speaker. The terminal controls the speaker and outputs the audio message at an appropriate volume. The input is the audio data sent to the terminal, and the output is the audio message played by the speaker.

[1005] Step 8:

[1006] A user can check the status of their home from outside using a dedicated smartphone app. The user operates the app and sends requests to the server. The input is the user's request, and the output is the request sent to the server.

[1007] Step 9:

[1008] The server communicates with the terminal and collects information on the situation near the entrance in real time. The server acquires data from cameras and sensors and analyzes the real-time images and situations. The input is a request from the user, and the output is the collected real-time data.

[1009] Step 10:

[1010] The server sends the collected real-time data back to the user's smartphone app, where the user can view the data. The input is the collected real-time data, and the output is the data sent to the user's smartphone app.

[1011] (Application example 1)

[1012] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1013] In conventional factories, it has sometimes been difficult to reliably identify visitors and employees and respond appropriately. As a result, there have been many issues with factory crime prevention measures and operational efficiency. For example, there is a need to strengthen security to prevent unauthorized visitors from entering the factory without permission, and to provide prompt and appropriate hospitality to employees. The purpose of the present invention is to solve these issues and improve security and operational efficiency within the factory.

[1014] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1015] In this invention, the server includes a means for identifying visitors using image analysis technology, a means for generating and outputting voice messages, and a means for controlling the voice output device according to a predetermined algorithm. This makes it possible to output a message such as "Welcome back" to registered users and a message such as "Hello. A representative will be there shortly, so please wait" to unregistered visitors. This makes it possible to strengthen crime prevention measures within the factory while providing appropriate hospitality to employees.

[1016] A "photography device" is a device installed near the entrance to capture images of visitors.

[1017] The "communication device" is a device for transmitting acquired image data to a server.

[1018] "Image analysis technology" is a technology that analyzes acquired image data and identifies visitors.

[1019] "User" is a term that refers to a pre-registered individual or employee.

[1020] "Unregistered Visitor" is a term used to refer to an individual or visitor who is not registered.

[1021] A "server" is a central processing unit for receiving and analyzing image data.

[1022] A "voice message" is a message in voice format that is generated and output based on the determination result of the server.

[1023] The "audio output device" is a device for outputting the generated audio message to the visitor.

[1024] The "predetermined algorithm" is a rule or calculation method that defines a certain procedure used by the server when controlling the audio output device.

[1025] As an embodiment of the present invention, a visitor / employee identification system for a factory will be described. This system is composed of a camera device, a communication device, a server, and an audio output device installed near the entrance.

[1026] System Configuration

[1027] 1. Imaging equipment

[1028] A camera (e.g., Logitech C920 HD Pro) installed near the entrance is responsible for capturing facial images of visitors and employees. The camera is controlled using the OpenCV library.

[1029] 2. Communications Equipment

[1030] The acquired facial image data is transmitted to a server via a communication device, using Internet Protocol to ensure secure transfer of data.

[1031] 3. Server

[1032] The server processes the image data using image analysis technology to identify visitors and employees. Specifically, it uses a pre-trained AI model (using the scikit-learn library) for facial recognition. Based on the identification results, the server generates an appropriate voice message.

[1033] 4. Audio Output Device

[1034] The voice output device outputs the generated voice messages to visitors and employees. This is done using the text_to_speech library.

[1035] Program processing explanation

[1036] Server Action:

[1037] The server first receives the image data sent from the camera. This image data is input into an AI model to identify whether the visitor is a registered employee or an unregistered visitor. Based on the identification result, a voice message such as "Welcome back" or "Hello. A representative will be there shortly, so please wait." is generated.

[1038] Audio output device handling:

[1039] The voice message received from the server is output through the voice output device. Specifically, the speak function is called to play the voice message from the speaker.

[1040] Specific examples

[1041] Example 1:

[1042] When a factory employee enters the entrance:

[1043] 1. The camera captures a facial image.

[1044] 2. The image data is sent to the server via the communication device.

[1045] 3. The server uses an AI model to identify the employee and generate a "Welcome back" message.

[1046] 4. The audio output device outputs the message.

[1047] Example 2:

[1048] When an unregistered visitor enters the entrance:

[1049] 1. The camera captures a facial image.

[1050] 2. The image data is sent to the server via the communication device.

[1051] 3. The server uses an AI model to identify the visitor as unregistered and generates a message saying, "Hello, please wait; a representative will be with you shortly."

[1052] 4. The audio output device outputs the message.

[1053] Example prompts to input to the generative AI model

[1054] "Generate a program to realize a system that identifies employees and visitors at the entrance of a factory and outputs a voice message saying "Welcome back" to employees and "Hello. A representative will be with you shortly" to visitors. Imagine an application that uses OpenCV and scikit-learn, working in conjunction with a camera and speaker."

[1055] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1056] Step 1:

[1057] The terminal (photographing device) captures the facial images of visitors who come near the entrance. The input is a real-time facial image, and the output is an image file in JPEG format or similar. This processing is performed using the OpenCV library.

[1058] Step 2:

[1059] The device sends the captured image data to a server over the Internet. The input is the captured image file, and the output is packets of transmitted data. The data is transferred securely using a communication protocol.

[1060] Step 3:

[1061] The server inputs the received image data into an AI model to identify the visitor. The input is the transferred image file, and the output is the identification result (e.g., registered employee, unregistered visitor). Scikit-learn is used as the AI ​​model, and image recognition is performed using a pre-trained model.

[1062] Step 4:

[1063] The server generates an appropriate voice message based on the identification results. The input is the identification results, and the output is a text message (e.g., "Welcome back," "Hello. A representative will be there shortly, so please wait"). If necessary, the generative AI model adjusts the details of the message.

[1064] Step 5:

[1065] The server then sends the generated voice message to the terminal's voice output device, with the input being the textual voice message and the output being packets of data to be transmitted, again using a communications protocol.

[1066] Step 6:

[1067] The terminal's audio output device converts the received voice message into speech and outputs it to the visitor through the speaker. The input is the text-based voice message, and the output is the actual voice message. This process is performed using the text_to_speech library.

[1068] Step 7:

[1069] A user (e.g., a factory manager) checks the status of their home from outside using a smartphone app. The input is a confirmation request from the user, and the output is real-time video data showing the situation near the entrance. The video is transferred from the server to the smartphone, and the user can check it using the app.

[1070] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1071] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[1072] System Configuration

[1073] The system mainly consists of the following components:

[1074] 1. Terminal: Includes a camera, microphone, speaker, and control device installed near the entrance.

[1075] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[1076] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[1077] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[1078] Program processing explanation

[1079] Below, the processing of the program for the entire system will be explained in natural language.

[1080] Image capture and transmission

[1081] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[1082] Facial Recognition and User Identification

[1083] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[1084] Emotion Recognition and Voice Message Generation

[1085] The server then uses an emotion engine to analyze the visitor's emotions from the received facial images and audio data, which can detect positive, negative, or neutral emotional states.

[1086] Based on the results of the identification and emotion analysis, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, a cheerful message such as "Welcome back, you have a lovely smile today!" is generated. If a negative emotion is detected, an encouraging message such as "Welcome back, you seem a little tired, are you okay?" is generated.

[1087] The server returns the generated voice message to the terminal, and the terminal outputs the received voice message from a speaker.

[1088] Check the status from outside

[1089] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[1090] Specific examples

[1091] Example 1: Registered user goes home

[1092] A user (Mr. Tanaka) enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses facial recognition technology to identify Mr. Tanaka and uses an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message "Welcome back, you have a lovely smile today!" from its speaker.

[1093] Example 2: Unregistered Visitor

[1094] A visitor (Mr. Sato) enters the front door. The device uses a camera to capture an image of Mr. Sato's face and sends it to the server. The server uses facial recognition technology to identify Mr. Sato as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[1095] Example 3: Checking the status while away from home

[1096] While Mr. Tanaka is out, he uses an app on his smartphone to check the situation at home. He sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitor's emotion recognition results, and sends them back to Mr. Tanaka. Mr. Tanaka can check the information in the app and send additional instructions to the server from the app if necessary.

[1097] Specific steps for person recognition and emotion analysis

[1098] 1. The device captures the visitor's facial image using a camera installed at the entrance and sends it to the server.

[1099] 2. The server uses facial recognition technology to identify visitors and determine whether they are registered users or unregistered visitors.

[1100] 3. The server analyzes the identified visitor's emotions using an emotion engine.

[1101] 4. The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[1102] 5. The server returns the generated voice message to the terminal, and the terminal outputs the voice from the speaker.

[1103] In this way, the present invention is a system that not only identifies visitors but also automatically responds appropriately according to their emotions, providing both crime prevention measures and hospitality at the same time.

[1104] The processing flow will be explained below.

[1105] Step 1:

[1106] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[1107] Step 2:

[1108] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[1109] Step 3:

[1110] The terminal transmits the captured image data to a server via the Internet.

[1111] Step 4:

[1112] The server inputs the received image data into an AI model (face recognition model).

[1113] Step 5:

[1114] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[1115] Step 6:

[1116] The server checks the identification result against an internal database to see if the corresponding user is registered.

[1117] Step 7:

[1118] The server determines the user identification result (registered user / unregistered user).

[1119] Step 8:

[1120] The server inputs the facial image data of the identified visitor into the emotion engine.

[1121] Step 9:

[1122] The server uses an emotion engine to analyze the visitor's emotions and detect positive, negative, or neutral emotional states.

[1123] Step 10:

[1124] The server generates an appropriate voice message based on the user identification and emotion analysis results. For example, if a positive emotion is detected in a registered user, the server generates a message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates a message such as "Welcome back, you seem a little tired, are you okay?"

[1125] Step 11:

[1126] The server returns the generated voice message to the terminal.

[1127] Step 12:

[1128] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[1129] Step 13:

[1130] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[1131] Step 14:

[1132] The server collects real-time data on the home situation (such as video streams and emotion recognition results) and sends it back to the user.

[1133] Step 15:

[1134] Users can check the requested real-time information on their smartphone app.

[1135] Step 16:

[1136] Users can enter additional instructions or messages through the app and send them to the server.

[1137] Step 17:

[1138] The server transfers the received instructions and messages to the terminal.

[1139] Step 18:

[1140] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[1141] Example 2

[1142] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1143] Conventional visitor identification systems are limited to identifying visitors through facial recognition and have the problem of being unable to respond to visitors with consideration for their emotional state. As a result, not only are hospitality towards visitors lacking, but they are also insufficient as a crime prevention measure. The purpose of this invention is to simultaneously strengthen hospitality and crime prevention measures by identifying the emotional state of visitors and automatically responding appropriately accordingly.

[1144] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1145] In this invention, the server includes means for analyzing emotions from received facial images, means for generating and outputting an appropriate voice message based on the analyzed emotions, and means for determining whether the visitor is a registered user or an unregistered visitor, thereby enabling a response that takes into account the visitor's emotional state.

[1146] A "camera" is a photographic device installed near the entrance to capture images of visitors.

[1147] A "terminal" is a control device including a camera, and is a device that detects visitors, captures images, and transmits image data.

[1148] The "server" is an information processing device that processes received image data, identifies visitors, analyzes their emotions, and generates voice messages.

[1149] "Image data" is digital data of the visitor's facial image captured by the camera.

[1150] "Image recognition technology" is a technology for analyzing image data and identifying specific people.

[1151] A "registered user" is a visitor whose face image data has been registered in advance in the system.

[1152] A "voice message" is a voice notification message that is output to a visitor.

[1153] The "emotion engine" is a software module for analyzing emotional states from facial images and voice data.

[1154] "Emotion analysis" is the process of analyzing a visitor's facial image and voice data to determine their emotional state.

[1155] "Positive emotions" refer to positive emotional states such as joy and satisfaction.

[1156] "Negative affect" refers to negative emotional states such as sadness or dissatisfaction.

[1157] A "Secure Data Transfer Protocol" is a communications method for safely and securely transferring data over the Internet.

[1158] "Hospitality" means treating visitors with a hospitable attitude.

[1159] "Crime prevention measures" are measures to prevent crime and fraud through visitor identification and emotion analysis.

[1160] MODE FOR CARRYING OUT THE INVENTION

[1161] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[1162] System Configuration

[1163] The system mainly consists of the following components:

[1164] 1. Terminal: A device equipped with a camera, microphone, speaker, and control unit installed near the entrance. The terminal captures the visitor's facial image and transmits the data to the server.

[1165] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[1166] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[1167] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[1168] Program processing

[1169] Below, the processing of the program for the entire system will be explained in natural language.

[1170] 1. Image capture and transmission

[1171] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[1172] 2. Facial Recognition and User Identification

[1173] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[1174] 3. Emotion Recognition and Voice Message Generation

[1175] Next, the server uses an emotion engine to analyze the visitor's emotions from the received facial images and voice data. The emotion engine can detect emotional states such as positive, negative, and neutral. Based on the identification and emotion analysis results, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, the server generates a cheerful message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates an encouraging message such as "Welcome back, you seem a little tired, are you okay?" The server then sends the generated voice message back to the terminal, and the terminal outputs the received voice message from its speaker.

[1176] 4. Check the status while on the go

[1177] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[1178] Specific example of system operation

[1179] Example 1: Registered user goes home

[1180] A user enters the front door. The device captures the user's facial image with a camera and sends it to the server. The server uses facial recognition technology to identify the user and an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message from its speaker: "Welcome back, you have a lovely smile today!"

[1181] Example 2: Unregistered Visitor

[1182] A visitor enters the front door. The device uses a camera to capture the visitor's facial image and sends it to the server. The server uses facial recognition technology to identify the visitor as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device then outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[1183] Example 3: Checking the status while away from home

[1184] While the user is out, they can check the status of their home using a smartphone app. The user sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitors' emotion recognition results, and sends them back to the user. The user can then check the information in the app and, if necessary, send additional instructions to the server.

[1185] Prompt Sentence Examples

[1186] A concrete example of a prompt might be:

[1187] "Facial image captured. Identification result is unregistered user. Emotion analysis results indicate the customer is in a negative emotional state."

[1188] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1189] Processing Step Description

[1190] Step 1: Visitor detection

[1191] The device detects human movement using a motion sensor installed at the entrance, and when the sensor responds, the device activates the camera.

[1192] Input: Motion sensor detection signal

[1193] Output: Camera activation signal

[1194] Specific operation: When the motion sensor is activated, the control device turns on the camera.

[1195] Step 2: Capture a face image

[1196] The device uses the activated camera to capture a facial image of the visitor.

[1197] Input: Camera status

[1198] Output: Face image data

[1199] What it does: The camera lens automatically adjusts focus and captures high-resolution images.

[1200] Step 3: Sending facial image data

[1201] The terminal transmits the captured facial image data to a server via the Internet.

[1202] Input: Face image data

[1203] Output: Image data sent to the server

[1204] What it does: The device encrypts the data and sends it securely to the server using the HTTPs protocol.

[1205] Step 4: Receiving facial image data

[1206] The server receives the face image data sent from the terminal.

[1207] Input: Image data from the device

[1208] Output: Received image data

[1209] Specific operation: Image data is temporarily stored in a database within the server for subsequent processing.

[1210] Step 5: Facial Recognition

[1211] The server inputs the received facial image data into an AI model to determine whether the visitor is a registered user or an unregistered visitor.

[1212] Input: Received image data

[1213] Output: Recognition results (registered users / unregistered visitors)

[1214] Specific operation: The server preprocesses (normalizes and resizes) the facial image, inputs it into the AI ​​model to perform facial recognition, and obtains a matching profile.

[1215] Step 6: Sentiment Analysis

[1216] The server uses an emotion engine to analyze emotions based on the image data and analysis results after facial recognition.

[1217] Input: Recognized image data

[1218] Output: Sentiment analysis result (positive / negative / neutral)

[1219] Specific operations: Extract facial feature points and analyze them to determine emotional state, and also analyze the voice spectrum if there are voice characteristics.

[1220] Step 7: Generate a voice message

[1221] The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[1222] Input: Recognition results, emotion analysis results

[1223] Output: The generated voice message

[1224] Specific operation: Using templates and analysis results, dynamically generate voice messages in text format and convert them into audio files.

[1225] Step 8: Send a voice message

[1226] The server transmits the generated voice message to the terminal.

[1227] Input: The generated voice message

[1228] Output: Audio message sent to the device

[1229] What it does: Compresses the voice message file and sends it to the device using a secure transfer protocol.

[1230] Step 9: Output a voice message

[1231] The terminal outputs the voice message received from the server from a speaker.

[1232] Input: Voice message sent to the device

[1233] Output: A voice message played through the speaker

[1234] Specific operation: The device decodes the received audio file and plays it on the internal speaker.

[1235] Step 10: Check the status on the go

[1236] Users can check the status of their home using a dedicated app.

[1237] Input: User request

[1238] Output: Real-time video and emotion recognition results

[1239] How it works: When a user sends a request from the app, the server collects the latest video data and emotion analysis results and immediately provides feedback to the user.

[1240] (Application example 2)

[1241] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1242] Conventional visitor response systems only use facial recognition technology to identify visitors, but do not take into account their emotional state. As a result, they are unable to determine the emotional state of the visitor and respond appropriately accordingly. Furthermore, there is a lack of technological means to improve hospitality for customers in brick-and-mortar stores.

[1243] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing facial image data to recognize the visitor's emotion, and means for generating and outputting an appropriate voice message based on the judgment result and emotion recognition result from the server. This makes it possible to provide individual responses according to the emotional state of the visitor, improving hospitality and providing optimal visitor service.

[1244] A "camera" is a device for capturing images and videos.

[1245] "Image data" is a digital representation of captured image information.

[1246] A "server" is a computer system that processes data and provides services over a network.

[1247] "Image recognition technology" is a technology that analyzes digital images and identifies specific objects or people.

[1248] A "visitor" refers to a person who visits a particular location.

[1249] A "registered user" is a specific person whose information is stored in advance in a database.

[1250] "Emotion recognition" is a technology that analyzes and recognizes a person's emotional state from image and audio information.

[1251] A "voice message" is a message for conveying information using voice.

[1252] The "judgment result" is the analysis result derived by the server using image recognition technology and emotion recognition technology.

[1253] The system for implementing this invention mainly captures images of visitors, identifies them based on the images, and then recognizes their emotions and outputs appropriate voice messages. This system is composed of a camera, a server, and a user terminal. Specific embodiments of the system are described below.

[1254] System Configuration

[1255] 1. Camera

[1256] The camera will be installed near the entrance and will capture facial images of visitors. The camera has high resolution and can take images in real time.

[1257] 2. Server

[1258] The server receives image data sent from the camera and identifies the visitor using image recognition technology. It then analyzes the visitor's emotions using emotion recognition technology (emotion engine). The server generates an appropriate voice message based on the identification and emotion analysis results and sends the voice message to the device.

[1259] 3. User Device

[1260] User terminals are devices such as smartphones, tablets, or smart glasses that can be used to check the system status and send instructions, including how to respond to visitors.

[1261] Program processing explanation

[1262] The server uses the following hardware and software:

[1263] Hardware

[1264] High-resolution cameras: Installed near the entrance to capture facial images of visitors.

[1265] Server: A powerful computer system for analyzing and processing data.

[1266] software

[1267] OpenCV: Used for image capture and processing.

[1268] EmotionRecognizer: A software module for analyzing emotions.

[1269] Server application: A dedicated application that receives image data, analyzes it, and generates voice messages.

[1270] The server receives image data captured by the camera and performs facial recognition using OpenCV. It then analyzes the visitor's emotions using EmotionRecognizer. Based on the analysis results, it generates an appropriate voice message and sends it back to the device to optimize visitor interaction.

[1271] Specific examples

[1272] Specific examples are given below:

[1273] Example 1: Response to a registered user's visit

[1274] 1. When a visitor approaches the entrance, a camera captures their facial image.

[1275] 2. The image data is sent to the server.

[1276] 3. The server uses facial recognition technology to identify the visitor as a registered user.

[1277] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as positive, for example.

[1278] 5. The server generates a voice message saying "Welcome back! Have a great day today!" and sends it to the device.

[1279] Example 2: How to handle unregistered visitors

[1280] 1. When a visitor approaches the entrance, a camera captures their facial image.

[1281] 2. The image data is sent to the server.

[1282] 3. The server uses facial recognition technology to identify the visitor as an unregistered person.

[1283] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as negative, for example.

[1284] 5. The server generates a voice message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it to the device.

[1285] Prompt Sentence Examples

[1286] "Taking a facial image as input, analyze the emotion as positive, negative, or neutral, and generate an appropriate message based on the analysis results."

[1287] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1288] Step 1:

[1289] A camera is installed near the entrance, and when a visitor approaches, the sensor detects their movement. The camera captures the visitor's facial image. The input here is the "visitor's facial image" and the output is the "captured visitor's facial image data."

[1290] Step 2:

[1291] The terminal sends the captured facial image data to a server via the Internet. The input here is the "captured facial image data of the visitor" and the output is the "facial image data sent to the server."

[1292] Step 3:

[1293] The server processes the received facial image data and identifies the visitor using image recognition technology. In this process, the server analyzes the facial image data using an AI model. The input is the "facial image data sent to the server" and the output is "information on the identified visitor."

[1294] Step 4:

[1295] Based on the identification result, the server refers to the database to determine whether the visitor is a registered user. At this time, it searches the database and obtains a response. The input is "identified visitor information" and the output is "the determination result of whether the user is registered or not."

[1296] Step 5:

[1297] The server uses an emotion engine based on the visitor's image data to perform emotion analysis. Emotion analysis uses a generative AI model to analyze emotions from facial expressions and other features. The input is the visitor's facial image data, and the output is the analyzed emotion data.

[1298] Step 6:

[1299] The server generates an appropriate voice message based on the results of the classification and emotion analysis. The generated voice message is based on a pre-defined prompt. The input is the "judgment result and emotion data," and the output is the "generated voice message."

[1300] Step 7:

[1301] The server sends the generated voice message to the terminal, and the terminal outputs the voice message through a speaker. Here, the input is the "generated voice message" and the output is the "voice output for the visitor."

[1302] Step 8:

[1303] A user can use a dedicated app to check the status of their home while away from home and send a request to the server. The input is a "request from the user" and the output is a "data request to the server."

[1304] Step 9:

[1305] The server collects data such as real-time video and emotion recognition results and returns them to the user. The input is an external data collection request, and the output is the data returned to the user.

[1306] Step 10:

[1307] Through the app, users can check the collected data and send additional instructions to the server if necessary. The input is "collected data" and the output is "information display and additional instructions to the user."

[1308] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1309] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1310] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1311] [Fourth embodiment]

[1312] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1313] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1314] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1315] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1316] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1317] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1318] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1319] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1320] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1321] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1322] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1323] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1324] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1325] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[1326] System Configuration

[1327] The system mainly consists of the following components:

[1328] 1. Terminal: This terminal is installed near the entrance and includes a camera, microphone, speaker, and control device. It captures the visitor's facial image and outputs a voice message.

[1329] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. Based on the results of the judgment, it generates a voice message and sends it back to the terminal.

[1330] 3. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home or send instructions while away from home.

[1331] Program processing explanation

[1332] Below, the processing of the program for the entire system will be explained in natural language.

[1333] Image capture and transmission

[1334] The terminal captures the visitor's facial image using a camera installed at the entrance, and transmits this image data to a server via the Internet.

[1335] Facial Recognition and User Identification

[1336] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[1337] Voice message output

[1338] The server generates a voice message based on the result of the judgment. Specifically, for registered users, it greets them with "Welcome back," and for unregistered visitors, it generates a warning message saying "Welcome. This area is being recorded." This voice message is then sent back to the terminal.

[1339] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[1340] Check the status from outside

[1341] Users can check the status of their homes from outside using a dedicated smartphone app, and can send requests to the server through this app.

[1342] The server collects information on the situation near the entrance in real time and sends the data back to the user, who can then check the information on their smartphone app.

[1343] Specific examples

[1344] Example 1: Registered user goes home

[1345] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[1346] Example 2: Unregistered Visitor

[1347] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[1348] Example 3: Checking the status while away from home

[1349] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[1350] The processing flow will be explained below.

[1351] Step 1:

[1352] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[1353] Step 2:

[1354] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[1355] Step 3:

[1356] The terminal transmits the captured image data to a server via the Internet.

[1357] Step 4:

[1358] The server inputs the received image data into an AI model (face recognition model).

[1359] Step 5:

[1360] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[1361] Step 6:

[1362] The server checks the identification result against an internal database to see if the corresponding user is registered.

[1363] Step 7:

[1364] The server determines the user identification result (registered user / unregistered user) and generates a voice message based on the result.

[1365] Step 8:

[1366] The server returns the generated voice message to the terminal.

[1367] Step 9:

[1368] The device outputs the received voice message from the speaker and responds appropriately to the visitor. Specifically, it greets the visitor with "Welcome back" if the visitor is a registered user, and warns the visitor that "Welcome. This area is being recorded" if the visitor is not registered.

[1369] Step 10:

[1370] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[1371] Step 11:

[1372] The server collects real-time data on the home situation (video streams and sensor information) and sends it back to the user.

[1373] Step 12:

[1374] Users can check the requested real-time information on their smartphone app.

[1375] Step 13:

[1376] Users can enter additional instructions or messages through the app and send them to the server.

[1377] Step 14:

[1378] The server transfers the received instructions and messages to the terminal.

[1379] Step 15:

[1380] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[1381] Example 1

[1382] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1383] While crime prevention measures are now required in homes and offices, it is also important to respond appropriately to visitors. However, existing systems do not adequately distinguish between registered and unregistered visitors and automate the response. In addition, there are limited ways to check the status of the home while away from home, and there is no way to check in real time, making it difficult to balance crime prevention and hospitality.

[1384] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1385] In this invention, the server includes means for capturing an image of the visitor using a camera installed near the entrance, means for transmitting the captured image data to a processing device, means for identifying the visitor using image analysis technology on the processing device, means for determining on the processing device whether the visitor is a registered user, and means for generating and outputting a voice message based on the determination result from the processing device. This makes it possible to automatically determine whether the visitor is a registered user or an unregistered visitor and output an appropriate voice message. Furthermore, the status of the home can be checked in real time from outside using a dedicated application, achieving both security and hospitality.

[1386] "Capture devices" are cameras or other image capture devices installed near entrances to capture images of visitors.

[1387] "Captured image data" refers to facial images and other video data of visitors captured by a camera.

[1388] A "processing device" is a device with computing resources such as a server, which analyzes image data, identifies visitors, and generates voice messages.

[1389] "Image analysis technology" is a technology that uses AI models and facial recognition software to identify visitors based on acquired image data.

[1390] A "registered user" refers to a person whose facial image data, etc. has been registered in the system in advance.

[1391] An "unregistered visitor" refers to a person whose facial image data or other data is not registered in the system.

[1392] A "voice message" is text data converted into voice data, and is the voice that the system plays to the visitor.

[1393] A "dedicated application" is software for smartphones or tablets that allows users to check the status of their home in real time while they are away from home.

[1394] "Real-time" means that data is collected and displayed close to the moment a visitor is near the entrance.

[1395] The present invention provides a system that distinguishes between registered users and unregistered visitors and outputs a voice message accordingly, thereby providing security measures and hospitality. Specific embodiments for carrying out the present invention will be described below.

[1396] System Configuration

[1397] The system mainly consists of the following components:

[1398] 1. Terminal: This includes a camera and audio output device installed near the entrance, as well as a control device that controls them. It is responsible for capturing facial images of visitors and outputting audio messages. Specifically, it uses a control device such as a Raspberry Pi, a Raspberry Pi camera module, and a speaker module.

[1399] 2. Server: A processing device that receives image data and identifies visitors using image analysis technology. Specifically, it performs facial recognition using TensorFlow and OpenCV libraries. It then generates a voice message based on the results of its judgment and sends it back to the device.

[1400] 3. User device: A device such as a smartphone or tablet that can be used to check the status of the home or office while away from home or to send instructions. Specifically, a dedicated application developed with Flutter is used.

[1401] Program processing explanation

[1402] The terminal uses a camera installed at the entrance to capture a facial image of the visitor. This image data is sent to a server via the Internet. The server inputs the received image data into an AI model and identifies the visitor using image analysis technology. This AI model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor. The server generates a voice message based on the judgment result. Specifically, if the visitor is a registered user, it greets them with "Welcome back," and if the visitor is an unregistered visitor, it generates a warning message saying "Welcome. This area is being recorded." This voice message is sent back to the terminal. The terminal then outputs the received voice message from its speaker and takes appropriate action against the visitor.

[1403] Users can check the status of their home from outside using a dedicated smartphone app. Through this app, they can send requests to the server. The server collects information about the situation near the entrance in real time and sends the data back to the user. The user can then check the information on the smartphone app.

[1404] Specific examples

[1405] Example of a registered user returning home

[1406] A user (Mr. Tanaka) returns home and enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses an AI model to identify Mr. Tanaka and determines that he is a registered user. Based on the server's judgment, the device outputs a voice message saying "Welcome home."

[1407] Unregistered visitor example

[1408] A visitor (Mr. Sato) enters the front door. The device captures Mr. Sato's facial image with its camera and sends it to the server. The server uses an AI model to identify Mr. Sato and determines that he is an unregistered visitor. Based on the server's judgment, the device outputs a warning message saying, "Welcome. This area is being recorded."

[1409] Example of checking the status from outside

[1410] While Tanaka is out, he checks the situation at home using the app on his smartphone. Tanaka sends a request from the app to the server. The server collects video footage of the area around the entrance in real time and sends it back to Tanaka. Tanaka then checks the situation at home using the app on his smartphone.

[1411] Example prompts for generative AI models

[1412] Example prompt sentence:

[1413] "Please generate a greeting message for the following visitor. The visitor is registered as Tanaka. ''"

[1414] Example of the resulting result:

[1415] "Welcome back, Tanaka-san. How was your day?"

[1416] Example prompt sentence:

[1417] "Generate a warning message when a visitor is not registered. ''"

[1418] Example of the resulting result:

[1419] "Welcome. We're recording here."

[1420] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1421] Step 1:

[1422] The terminal captures the visitor's facial image using a camera installed at the entrance. When a visitor enters the entrance, the sensor reacts and the camera captures the facial image. The input is the visitor's facial image, and the output is the captured image data.

[1423] Step 2:

[1424] The terminal transmits the image data captured by the imaging device to a server via the Internet. The HTTP protocol is used for transmission, and the image data is encoded and sent to the server. The input is the captured image data, and the output is the image data sent to the server.

[1425] Step 3:

[1426] The server temporarily stores the received image data and prepares it for use as input to the AI ​​model. At this time, the image data is decoded and stored in memory. The input is the image data sent to the server, and the output is image data converted into an analyzable format.

[1427] Step 4:

[1428] The server inputs image data into an AI model to perform facial recognition. The AI ​​model is trained based on facial image data of pre-registered users. Libraries such as TensorFlow and OpenCV are used for the identification process. The input is image data converted into an analyzable format, and the output is the identification result of whether the visitor is a registered user or an unregistered visitor.

[1429] Step 5:

[1430] The server generates a voice message based on the identification result. For registered users, it generates a text message such as "Welcome back," and for unregistered visitors, it converts the text message into voice data using the Google Text-to-Speech API. The input is the identification result, and the output is the generated voice data.

[1431] Step 6:

[1432] The server sends the generated audio data to the terminal, again using the HTTP protocol. The input is the generated audio data, and the output is the audio data sent to the terminal.

[1433] Step 7:

[1434] The terminal plays the received audio data to the visitor through the speaker. The terminal controls the speaker and outputs the audio message at an appropriate volume. The input is the audio data sent to the terminal, and the output is the audio message played by the speaker.

[1435] Step 8:

[1436] A user can check the status of their home from outside using a dedicated smartphone app. The user operates the app and sends requests to the server. The input is the user's request, and the output is the request sent to the server.

[1437] Step 9:

[1438] The server communicates with the terminal and collects information on the situation near the entrance in real time. The server acquires data from cameras and sensors and analyzes the real-time images and situations. The input is a request from the user, and the output is the collected real-time data.

[1439] Step 10:

[1440] The server sends the collected real-time data back to the user's smartphone app, where the user can view the data. The input is the collected real-time data, and the output is the data sent to the user's smartphone app.

[1441] (Application example 1)

[1442] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1443] In conventional factories, it has sometimes been difficult to reliably identify visitors and employees and respond appropriately. As a result, there have been many issues with factory crime prevention measures and operational efficiency. For example, there is a need to strengthen security to prevent unauthorized visitors from entering the factory without permission, and to provide prompt and appropriate hospitality to employees. The purpose of the present invention is to solve these issues and improve security and operational efficiency within the factory.

[1444] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1445] In this invention, the server includes a means for identifying visitors using image analysis technology, a means for generating and outputting voice messages, and a means for controlling the voice output device according to a predetermined algorithm. This makes it possible to output a message such as "Welcome back" to registered users and a message such as "Hello. A representative will be there shortly, so please wait" to unregistered visitors. This makes it possible to strengthen crime prevention measures within the factory while providing appropriate hospitality to employees.

[1446] A "photography device" is a device installed near the entrance to capture images of visitors.

[1447] The "communication device" is a device for transmitting acquired image data to a server.

[1448] "Image analysis technology" is a technology that analyzes acquired image data and identifies visitors.

[1449] "User" is a term that refers to a pre-registered individual or employee.

[1450] "Unregistered Visitor" is a term used to refer to an individual or visitor who is not registered.

[1451] A "server" is a central processing unit for receiving and analyzing image data.

[1452] A "voice message" is a message in voice format that is generated and output based on the determination result of the server.

[1453] The "audio output device" is a device for outputting the generated audio message to the visitor.

[1454] The "predetermined algorithm" is a rule or calculation method that defines a certain procedure used by the server when controlling the audio output device.

[1455] As an embodiment of the present invention, a visitor / employee identification system for a factory will be described. This system is composed of a camera device, a communication device, a server, and an audio output device installed near the entrance.

[1456] System Configuration

[1457] 1. Imaging equipment

[1458] A camera (e.g., Logitech C920 HD Pro) installed near the entrance is responsible for capturing facial images of visitors and employees. The camera is controlled using the OpenCV library.

[1459] 2. Communications Equipment

[1460] The acquired facial image data is transmitted to a server via a communication device, using Internet Protocol to ensure secure transfer of data.

[1461] 3. Server

[1462] The server processes the image data using image analysis technology to identify visitors and employees. Specifically, it uses a pre-trained AI model (using the scikit-learn library) for facial recognition. Based on the identification results, the server generates an appropriate voice message.

[1463] 4. Audio Output Device

[1464] The voice output device outputs the generated voice messages to visitors and employees. This is done using the text_to_speech library.

[1465] Program processing explanation

[1466] Server Action:

[1467] The server first receives the image data sent from the camera. This image data is input into an AI model to identify whether the visitor is a registered employee or an unregistered visitor. Based on the identification result, a voice message such as "Welcome back" or "Hello. A representative will be there shortly, so please wait." is generated.

[1468] Audio output device handling:

[1469] The voice message received from the server is output through the voice output device. Specifically, the speak function is called to play the voice message from the speaker.

[1470] Specific examples

[1471] Example 1:

[1472] When a factory employee enters the entrance:

[1473] 1. The camera captures a facial image.

[1474] 2. The image data is sent to the server via the communication device.

[1475] 3. The server uses an AI model to identify the employee and generate a "Welcome back" message.

[1476] 4. The audio output device outputs the message.

[1477] Example 2:

[1478] When an unregistered visitor enters the entrance:

[1479] 1. The camera captures a facial image.

[1480] 2. The image data is sent to the server via the communication device.

[1481] 3. The server uses an AI model to identify the visitor as unregistered and generates a message saying, "Hello, please wait; a representative will be with you shortly."

[1482] 4. The audio output device outputs the message.

[1483] Example prompts to input to the generative AI model

[1484] "Generate a program to realize a system that identifies employees and visitors at the entrance of a factory and outputs a voice message saying "Welcome back" to employees and "Hello. A representative will be with you shortly" to visitors. Imagine an application that uses OpenCV and scikit-learn, working in conjunction with a camera and speaker."

[1485] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1486] Step 1:

[1487] The terminal (photographing device) captures the facial images of visitors who come near the entrance. The input is a real-time facial image, and the output is an image file in JPEG format or similar. This processing is performed using the OpenCV library.

[1488] Step 2:

[1489] The device sends the captured image data to a server over the Internet. The input is the captured image file, and the output is packets of transmitted data. The data is transferred securely using a communication protocol.

[1490] Step 3:

[1491] The server inputs the received image data into an AI model to identify the visitor. The input is the transferred image file, and the output is the identification result (e.g., registered employee, unregistered visitor). Scikit-learn is used as the AI ​​model, and image recognition is performed using a pre-trained model.

[1492] Step 4:

[1493] The server generates an appropriate voice message based on the identification results. The input is the identification results, and the output is a text message (e.g., "Welcome back," "Hello. A representative will be there shortly, so please wait"). If necessary, the generative AI model adjusts the details of the message.

[1494] Step 5:

[1495] The server then sends the generated voice message to the terminal's voice output device, with the input being the textual voice message and the output being packets of data to be transmitted, again using a communications protocol.

[1496] Step 6:

[1497] The terminal's audio output device converts the received voice message into speech and outputs it to the visitor through the speaker. The input is the text-based voice message, and the output is the actual voice message. This process is performed using the text_to_speech library.

[1498] Step 7:

[1499] A user (e.g., a factory manager) checks the status of their home from outside using a smartphone app. The input is a confirmation request from the user, and the output is real-time video data showing the situation near the entrance. The video is transferred from the server to the smartphone, and the user can check it using the app.

[1500] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1501] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[1502] System Configuration

[1503] The system mainly consists of the following components:

[1504] 1. Terminal: Includes a camera, microphone, speaker, and control device installed near the entrance.

[1505] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[1506] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[1507] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[1508] Program processing explanation

[1509] Below, the processing of the program for the entire system will be explained in natural language.

[1510] Image capture and transmission

[1511] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[1512] Facial Recognition and User Identification

[1513] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[1514] Emotion Recognition and Voice Message Generation

[1515] The server then uses an emotion engine to analyze the visitor's emotions from the received facial images and audio data, which can detect positive, negative, or neutral emotional states.

[1516] Based on the results of the identification and emotion analysis, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, a cheerful message such as "Welcome back, you have a lovely smile today!" is generated. If a negative emotion is detected, an encouraging message such as "Welcome back, you seem a little tired, are you okay?" is generated.

[1517] The server returns the generated voice message to the terminal, and the terminal outputs the received voice message from a speaker.

[1518] Check the status from outside

[1519] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[1520] Specific examples

[1521] Example 1: Registered user goes home

[1522] A user (Mr. Tanaka) enters the front door. The device captures an image of Mr. Tanaka's face with its camera and sends it to the server. The server uses facial recognition technology to identify Mr. Tanaka and uses an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message "Welcome back, you have a lovely smile today!" from its speaker.

[1523] Example 2: Unregistered Visitor

[1524] A visitor (Mr. Sato) enters the front door. The device uses a camera to capture an image of Mr. Sato's face and sends it to the server. The server uses facial recognition technology to identify Mr. Sato as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[1525] Example 3: Checking the status while away from home

[1526] While Mr. Tanaka is out, he uses an app on his smartphone to check the situation at home. He sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitor's emotion recognition results, and sends them back to Mr. Tanaka. Mr. Tanaka can check the information in the app and send additional instructions to the server from the app if necessary.

[1527] Specific steps for person recognition and emotion analysis

[1528] 1. The device captures the visitor's facial image using a camera installed at the entrance and sends it to the server.

[1529] 2. The server uses facial recognition technology to identify visitors and determine whether they are registered users or unregistered visitors.

[1530] 3. The server analyzes the identified visitor's emotions using an emotion engine.

[1531] 4. The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[1532] 5. The server returns the generated voice message to the terminal, and the terminal outputs the voice from the speaker.

[1533] In this way, the present invention is a system that not only identifies visitors but also automatically responds appropriately according to their emotions, providing both crime prevention measures and hospitality at the same time.

[1534] The processing flow will be explained below.

[1535] Step 1:

[1536] The device uses a sensor to detect when a person approaches the entrance and activates the camera.

[1537] Step 2:

[1538] The terminal captures the visitor's facial image with a camera and stores the image data in its internal memory.

[1539] Step 3:

[1540] The terminal transmits the captured image data to a server via the Internet.

[1541] Step 4:

[1542] The server inputs the received image data into an AI model (face recognition model).

[1543] Step 5:

[1544] The server uses an AI model to analyze the visitor's face and obtains the user ID or "unregistered" as the identification result.

[1545] Step 6:

[1546] The server checks the identification result against an internal database to see if the corresponding user is registered.

[1547] Step 7:

[1548] The server determines the user identification result (registered user / unregistered user).

[1549] Step 8:

[1550] The server inputs the facial image data of the identified visitor into the emotion engine.

[1551] Step 9:

[1552] The server uses an emotion engine to analyze the visitor's emotions and detect positive, negative, or neutral emotional states.

[1553] Step 10:

[1554] The server generates an appropriate voice message based on the user identification and emotion analysis results. For example, if a positive emotion is detected in a registered user, the server generates a message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates a message such as "Welcome back, you seem a little tired, are you okay?"

[1555] Step 11:

[1556] The server returns the generated voice message to the terminal.

[1557] Step 12:

[1558] The terminal outputs the received voice message from the speaker and responds appropriately to the visitor.

[1559] Step 13:

[1560] When a user is away from home, they launch a dedicated app on their smartphone and, when they want to check the situation at home, they send a request to the server through the app.

[1561] Step 14:

[1562] The server collects real-time data on the home situation (such as video streams and emotion recognition results) and sends it back to the user.

[1563] Step 15:

[1564] Users can check the requested real-time information on their smartphone app.

[1565] Step 16:

[1566] Users can enter additional instructions or messages through the app and send them to the server.

[1567] Step 17:

[1568] The server transfers the received instructions and messages to the terminal.

[1569] Step 18:

[1570] The device outputs transferred instructions and messages as voice through a speaker, supporting the exchange of messages between family members.

[1571] Example 2

[1572] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1573] Conventional visitor identification systems are limited to identifying visitors through facial recognition and have the problem of being unable to respond to visitors with consideration for their emotional state. As a result, not only are hospitality towards visitors lacking, but they are also insufficient as a crime prevention measure. The purpose of this invention is to simultaneously strengthen hospitality and crime prevention measures by identifying the emotional state of visitors and automatically responding appropriately accordingly.

[1574] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1575] In this invention, the server includes means for analyzing emotions from received facial images, means for generating and outputting an appropriate voice message based on the analyzed emotions, and means for determining whether the visitor is a registered user or an unregistered visitor, thereby enabling a response that takes into account the visitor's emotional state.

[1576] A "camera" is a photographic device installed near the entrance to capture images of visitors.

[1577] A "terminal" is a control device including a camera, and is a device that detects visitors, captures images, and transmits image data.

[1578] The "server" is an information processing device that processes received image data, identifies visitors, analyzes their emotions, and generates voice messages.

[1579] "Image data" is digital data of the visitor's facial image captured by the camera.

[1580] "Image recognition technology" is a technology for analyzing image data and identifying specific people.

[1581] A "registered user" is a visitor whose face image data has been registered in advance in the system.

[1582] A "voice message" is a voice notification message that is output to a visitor.

[1583] The "emotion engine" is a software module for analyzing emotional states from facial images and voice data.

[1584] "Emotion analysis" is the process of analyzing a visitor's facial image and voice data to determine their emotional state.

[1585] "Positive emotions" refer to positive emotional states such as joy and satisfaction.

[1586] "Negative affect" refers to negative emotional states such as sadness or dissatisfaction.

[1587] A "Secure Data Transfer Protocol" is a communications method for safely and securely transferring data over the Internet.

[1588] "Hospitality" means treating visitors with a hospitable attitude.

[1589] "Crime prevention measures" are measures to prevent crime and fraud through visitor identification and emotion analysis.

[1590] MODE FOR CARRYING OUT THE INVENTION

[1591] The present invention is a system that captures a facial image of a visitor, analyzes the image to identify the visitor, recognizes the visitor's emotions using an emotion engine, and generates and outputs an appropriate voice message. Specific embodiments for implementing the present invention will be described below.

[1592] System Configuration

[1593] The system mainly consists of the following components:

[1594] 1. Terminal: A device equipped with a camera, microphone, speaker, and control unit installed near the entrance. The terminal captures the visitor's facial image and transmits the data to the server.

[1595] 2. Server: A processing device that receives image data and identifies visitors using image recognition technology. It uses an emotion engine to analyze emotions and generate a voice message based on the judgment results.

[1596] 3. Emotion engine: A software module for analyzing emotions from the user's facial expressions and voice.

[1597] 4. User terminal: A device such as a smartphone or tablet that can be used to check the status of the home and send instructions while away from home.

[1598] Program processing

[1599] Below, the processing of the program for the entire system will be explained in natural language.

[1600] 1. Image capture and transmission

[1601] The device uses a sensor to detect when a person approaches the entrance and activates the camera, which captures the visitor's facial image and sends the image data to a server via the Internet.

[1602] 2. Facial Recognition and User Identification

[1603] The server inputs the received image data into an AI model and identifies the visitor using image recognition technology. The AI ​​model is trained based on facial image data of pre-registered users. The identification result determines whether the visitor is a registered user or an unregistered visitor.

[1604] 3. Emotion Recognition and Voice Message Generation

[1605] Next, the server uses an emotion engine to analyze the visitor's emotions from the received facial images and voice data. The emotion engine can detect emotional states such as positive, negative, and neutral. Based on the identification and emotion analysis results, the server generates an appropriate voice message. For example, if the user is registered and a positive emotion is detected, the server generates a cheerful message such as "Welcome back, you have a lovely smile today!". If a negative emotion is detected, the server generates an encouraging message such as "Welcome back, you seem a little tired, are you okay?" The server then sends the generated voice message back to the terminal, and the terminal outputs the received voice message from its speaker.

[1606] 4. Check the status while on the go

[1607] Users can check the status of their home from outside by launching a dedicated smartphone app. The app sends a request to the server, which then collects the corresponding data (such as real-time video and emotion recognition results) and returns it to the user.

[1608] Specific example of system operation

[1609] Example 1: Registered user goes home

[1610] A user enters the front door. The device captures the user's facial image with a camera and sends it to the server. The server uses facial recognition technology to identify the user and an emotion engine to detect positive emotions from the facial image. The server generates a voice message saying, "Welcome back, you have a lovely smile today!" and sends it back to the device. The device then outputs the message from its speaker: "Welcome back, you have a lovely smile today!"

[1611] Example 2: Unregistered Visitor

[1612] A visitor enters the front door. The device uses a camera to capture the visitor's facial image and sends it to the server. The server uses facial recognition technology to identify the visitor as an unregistered visitor and uses an emotion engine to detect negative emotions from the facial image. The server generates a warning message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it back to the device. The device then outputs a voice message from its speaker saying, "Welcome. This section is being recorded. Is there anything I can help you with?"

[1613] Example 3: Checking the status while away from home

[1614] While the user is out, they can check the status of their home using a smartphone app. The user sends a request from the app to the server. The server collects real-time video footage of the area around the entrance and the visitors' emotion recognition results, and sends them back to the user. The user can then check the information in the app and, if necessary, send additional instructions to the server.

[1615] Prompt Sentence Examples

[1616] A concrete example of a prompt might be:

[1617] "Facial image captured. Identification result is unregistered user. Emotion analysis results indicate the customer is in a negative emotional state."

[1618] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1619] Processing Step Description

[1620] Step 1: Visitor detection

[1621] The device detects human movement using a motion sensor installed at the entrance, and when the sensor responds, the device activates the camera.

[1622] Input: Motion sensor detection signal

[1623] Output: Camera activation signal

[1624] Specific operation: When the motion sensor is activated, the control device turns on the camera.

[1625] Step 2: Capture a face image

[1626] The device uses the activated camera to capture a facial image of the visitor.

[1627] Input: Camera status

[1628] Output: Face image data

[1629] What it does: The camera lens automatically adjusts focus and captures high-resolution images.

[1630] Step 3: Sending facial image data

[1631] The terminal transmits the captured facial image data to a server via the Internet.

[1632] Input: Face image data

[1633] Output: Image data sent to the server

[1634] What it does: The device encrypts the data and sends it securely to the server using the HTTPs protocol.

[1635] Step 4: Receiving facial image data

[1636] The server receives the face image data sent from the terminal.

[1637] Input: Image data from the device

[1638] Output: Received image data

[1639] Specific operation: Image data is temporarily stored in a database within the server for subsequent processing.

[1640] Step 5: Facial Recognition

[1641] The server inputs the received facial image data into an AI model to determine whether the visitor is a registered user or an unregistered visitor.

[1642] Input: Received image data

[1643] Output: Recognition results (registered users / unregistered visitors)

[1644] Specific operation: The server preprocesses (normalizes and resizes) the facial image, inputs it into the AI ​​model to perform facial recognition, and obtains a matching profile.

[1645] Step 6: Sentiment Analysis

[1646] The server uses an emotion engine to analyze emotions based on the image data and analysis results after facial recognition.

[1647] Input: Recognized image data

[1648] Output: Sentiment analysis result (positive / negative / neutral)

[1649] Specific operations: Extract facial feature points and analyze them to determine emotional state, and also analyze the voice spectrum if there are voice characteristics.

[1650] Step 7: Generate a voice message

[1651] The server generates an appropriate voice message based on the results of the identification and emotion analysis.

[1652] Input: Recognition results, emotion analysis results

[1653] Output: The generated voice message

[1654] Specific operation: Using templates and analysis results, dynamically generate voice messages in text format and convert them into audio files.

[1655] Step 8: Send a voice message

[1656] The server transmits the generated voice message to the terminal.

[1657] Input: The generated voice message

[1658] Output: Audio message sent to the device

[1659] What it does: Compresses the voice message file and sends it to the device using a secure transfer protocol.

[1660] Step 9: Output a voice message

[1661] The terminal outputs the voice message received from the server from a speaker.

[1662] Input: Voice message sent to the device

[1663] Output: A voice message played through the speaker

[1664] Specific operation: The device decodes the received audio file and plays it on the internal speaker.

[1665] Step 10: Check the status on the go

[1666] Users can check the status of their home using a dedicated app.

[1667] Input: User request

[1668] Output: Real-time video and emotion recognition results

[1669] How it works: When a user sends a request from the app, the server collects the latest video data and emotion analysis results and immediately provides feedback to the user.

[1670] (Application example 2)

[1671] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1672] Conventional visitor response systems only use facial recognition technology to identify visitors, but do not take into account their emotional state. As a result, they are unable to determine the emotional state of the visitor and respond appropriately accordingly. Furthermore, there is a lack of technological means to improve hospitality for customers in brick-and-mortar stores.

[1673] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing facial image data to recognize the visitor's emotion, and means for generating and outputting an appropriate voice message based on the judgment result and emotion recognition result from the server. This makes it possible to provide individual responses according to the emotional state of the visitor, improving hospitality and providing optimal visitor service.

[1674] A "camera" is a device for capturing images and videos.

[1675] "Image data" is a digital representation of captured image information.

[1676] A "server" is a computer system that processes data and provides services over a network.

[1677] "Image recognition technology" is a technology that analyzes digital images and identifies specific objects or people.

[1678] A "visitor" refers to a person who visits a particular location.

[1679] A "registered user" is a specific person whose information is stored in advance in a database.

[1680] "Emotion recognition" is a technology that analyzes and recognizes a person's emotional state from image and audio information.

[1681] A "voice message" is a message for conveying information using voice.

[1682] The "judgment result" is the analysis result derived by the server using image recognition technology and emotion recognition technology.

[1683] The system for implementing this invention mainly captures images of visitors, identifies them based on the images, and then recognizes their emotions and outputs appropriate voice messages. This system is composed of a camera, a server, and a user terminal. Specific embodiments of the system are described below.

[1684] System Configuration

[1685] 1. Camera

[1686] The camera will be installed near the entrance and will capture facial images of visitors. The camera has high resolution and can take images in real time.

[1687] 2. Server

[1688] The server receives image data sent from the camera and identifies the visitor using image recognition technology. It then analyzes the visitor's emotions using emotion recognition technology (emotion engine). The server generates an appropriate voice message based on the identification and emotion analysis results and sends the voice message to the device.

[1689] 3. User Device

[1690] User terminals are devices such as smartphones, tablets, or smart glasses that can be used to check the system status and send instructions, including how to respond to visitors.

[1691] Program processing explanation

[1692] The server uses the following hardware and software:

[1693] Hardware

[1694] High-resolution cameras: Installed near the entrance to capture facial images of visitors.

[1695] Server: A powerful computer system for analyzing and processing data.

[1696] software

[1697] OpenCV: Used for image capture and processing.

[1698] EmotionRecognizer: A software module for analyzing emotions.

[1699] Server application: A dedicated application that receives image data, analyzes it, and generates voice messages.

[1700] The server receives image data captured by the camera and performs facial recognition using OpenCV. It then analyzes the visitor's emotions using EmotionRecognizer. Based on the analysis results, it generates an appropriate voice message and sends it back to the device to optimize visitor interaction.

[1701] Specific examples

[1702] Specific examples are given below:

[1703] Example 1: Response to a registered user's visit

[1704] 1. When a visitor approaches the entrance, a camera captures their facial image.

[1705] 2. The image data is sent to the server.

[1706] 3. The server uses facial recognition technology to identify the visitor as a registered user.

[1707] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as positive, for example.

[1708] 5. The server generates a voice message saying "Welcome back! Have a great day today!" and sends it to the device.

[1709] Example 2: How to handle unregistered visitors

[1710] 1. When a visitor approaches the entrance, a camera captures their facial image.

[1711] 2. The image data is sent to the server.

[1712] 3. The server uses facial recognition technology to identify the visitor as an unregistered person.

[1713] 4. Then, the emotion engine is used to analyze the visitor's emotion, which is identified as negative, for example.

[1714] 5. The server generates a voice message saying, "Welcome. This section is being recorded. Is there anything I can help you with?" and sends it to the device.

[1715] Prompt Sentence Examples

[1716] "Taking a facial image as input, analyze the emotion as positive, negative, or neutral, and generate an appropriate message based on the analysis results."

[1717] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1718] Step 1:

[1719] A camera is installed near the entrance, and when a visitor approaches, the sensor detects their movement. The camera captures the visitor's facial image. The input here is the "visitor's facial image" and the output is the "captured visitor's facial image data."

[1720] Step 2:

[1721] The terminal sends the captured facial image data to a server via the Internet. The input here is the "captured facial image data of the visitor" and the output is the "facial image data sent to the server."

[1722] Step 3:

[1723] The server processes the received facial image data and identifies the visitor using image recognition technology. In this process, the server analyzes the facial image data using an AI model. The input is the "facial image data sent to the server" and the output is "information on the identified visitor."

[1724] Step 4:

[1725] Based on the identification result, the server refers to the database to determine whether the visitor is a registered user. At this time, it searches the database and obtains a response. The input is "identified visitor information" and the output is "the determination result of whether the user is registered or not."

[1726] Step 5:

[1727] The server uses an emotion engine based on the visitor's image data to perform emotion analysis. Emotion analysis uses a generative AI model to analyze emotions from facial expressions and other features. The input is the visitor's facial image data, and the output is the analyzed emotion data.

[1728] Step 6:

[1729] The server generates an appropriate voice message based on the results of the classification and emotion analysis. The generated voice message is based on a pre-defined prompt. The input is the "judgment result and emotion data," and the output is the "generated voice message."

[1730] Step 7:

[1731] The server sends the generated voice message to the terminal, and the terminal outputs the voice message through a speaker. Here, the input is the "generated voice message" and the output is the "voice output for the visitor."

[1732] Step 8:

[1733] A user can use a dedicated app to check the status of their home while away from home and send a request to the server. The input is a "request from the user" and the output is a "data request to the server."

[1734] Step 9:

[1735] The server collects data such as real-time video and emotion recognition results and returns them to the user. The input is an external data collection request, and the output is the data returned to the user.

[1736] Step 10:

[1737] Through the app, users can check the collected data and send additional instructions to the server if necessary. The input is "collected data" and the output is "information display and additional instructions to the user."

[1738] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1739] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1740] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1741] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1742] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1743] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1744] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1745] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1746] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1747] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1748] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1749] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1750] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1751] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1752] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1753] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1754] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1755] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1756] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1757] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1758] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1759] The following is further disclosed regarding the above embodiment.

[1760] (Claim 1)

[1761] means for capturing images of visitors using a camera located near the entrance;

[1762] means for transmitting the captured image data to a server;

[1763] A means for identifying visitors using image recognition technology on the server;

[1764] means on the server for determining whether the visitor is a registered user;

[1765] means for generating and outputting a voice message based on the determination result from the server;

[1766] A system including:

[1767] (Claim 2)

[1768] 10. The system of claim 1, wherein the system analyzes images captured by a camera near the entrance and identifies visitors using facial recognition technology.

[1769] (Claim 3)

[1770] The system according to claim 2, wherein a voice message saying "Welcome back" is output for a registered user, and a warning message saying "Welcome. This page is being recorded" is output for an unregistered visitor.

[1771] "Example 1"

[1772] (Claim 1)

[1773] a means for capturing an image of a visitor using a photographing device located near the entrance;

[1774] means for transmitting the acquired image data to a processing device;

[1775] means for identifying visitors using image analysis techniques on the processing device;

[1776] means for determining on the processing device whether the visitor is a registered user;

[1777] means for generating and outputting a voice message based on the determination result from the processing device;

[1778] A system including:

[1779] (Claim 2)

[1780] The system of claim 1 analyzes images captured by a camera near the entrance and identifies visitors using facial recognition technology.

[1781] (Claim 3)

[1782] The system according to claim 2, wherein the system outputs a voice message saying "Welcome back" to registered users, and outputs a warning message saying "Welcome. This is being recorded" to unregistered visitors.

[1783] "Application Example 1"

[1784] (Claim 1)

[1785] a means for capturing an image of a visitor using a photographing device located near the entrance;

[1786] means for transmitting the acquired image data to a server using a communication device;

[1787] A means for identifying visitors using image analysis technology on the server;

[1788] means on the server for determining whether the visitor is a registered user;

[1789] means for generating and outputting a voice message based on the determination result from the server;

[1790] means for controlling an audio output device in accordance with a predetermined algorithm;

[1791] A system including:

[1792] (Claim 2)

[1793] The system of claim 1 analyzes images captured by a camera near the entrance and identifies visitors using facial recognition technology.

[1794] (Claim 3)

[1795] The system of claim 1 outputs a voice message saying "Welcome back" for registered users, and outputs a guidance message saying "Hello. A representative will be with you shortly, so please wait" for unregistered visitors.

[1796] "Example 2: Combining Emotion Engines"

[1797] (Claim 1)

[1798] means for capturing images of visitors using a camera located near the entrance;

[1799] means for transmitting the captured image data to a server;

[1800] A means for identifying visitors using image recognition technology on the server;

[1801] means on the server for determining whether the visitor is a registered user;

[1802] A means for analyzing emotions from the received facial image on the server;

[1803] means for generating and outputting an appropriate voice message based on the analyzed emotion;

[1804] A system including:

[1805] (Claim 2)

[1806] 10. The system of claim 1, wherein the system analyzes images captured by a camera near the entrance and identifies visitors using facial recognition technology.

[1807] (Claim 3)

[1808] The system of claim 1 outputs a voice message saying "Welcome back, you have a lovely smile today!" when a positive emotion is detected in the case of a registered user, and outputs a warning message saying "Welcome. This page is being recorded. Is there anything I can help you with?" when a negative emotion is detected in the case of an unregistered visitor.

[1809] "Application example 2 when combining emotion engines"

[1810] (Claim 1)

[1811] means for capturing images of visitors using a camera located near the entrance;

[1812] means for transmitting the captured image data to a server;

[1813] A means for identifying visitors using image recognition technology on the server;

[1814] means on the server for determining whether the visitor is a registered user;

[1815] means for analyzing the received facial image data on the server to recognize the visitor's emotions;

[1816] means for generating and outputting an appropriate voice message based on the judgment result and emotion recognition result from the server;

[1817] A system including:

[1818] (Claim 2)

[1819] 10. The system of claim 1, wherein the system analyzes images captured by a camera near the entrance and identifies visitors using facial recognition technology.

[1820] (Claim 3)

[1821] 2. The system according to claim 1, wherein an appropriate voice message according to the visitor's emotions is output in the case of a registered user, and a warning message is output in the case of an unregistered visitor. [Explanation of symbols]

[1822] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for capturing images of visitors using a camera located near the entrance; means for transmitting the captured image data to a server; A means for identifying visitors using image recognition technology on the server; means on the server for determining whether the visitor is a registered user; means for generating and outputting a voice message based on the determination result from the server; A system including:

2. The system of claim 1 , wherein the system analyzes images captured by a camera near the entrance and identifies visitors using facial recognition technology.

3. 3. The system according to claim 2, wherein a voice message saying "Welcome back" is output to a registered user, and a warning message saying "Welcome. This page is being recorded" is output to an unregistered visitor.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A