System

The system addresses communication challenges in changing work environments by using facial recognition, generative AI, and automated responses to enhance efficiency and satisfaction.

JP2026033982APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137103
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Recent changes in the work environment, such as fewer employees in the office, mask-wearing, and telecommuting, hinder smooth communication and increase the workload by making facial recognition difficult and requiring after-hours responses to inquiries, negatively impacting work efficiency and employee satisfaction.

Method used

A system integrating facial recognition, data analysis using generative AI models, voice recognition, automatic responses after work, and data transmission/reception to identify colleagues, provide profile information, and handle inquiries efficiently.

Benefits of technology

Facilitates smooth communication and improved work efficiency by enabling easy identification of colleagues, reducing workload through voice commands and automated responses, and ensuring timely information delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026033982000001_ABST
    Figure 2026033982000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a face recognition means, a AI analysis means by a generation sound model, an operation means by sound recognition, an automatic dealing means after the end of work, and a data-transmitting / receiving means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Recent changes in the work environment have led to an increasing number of cases where smooth communication in business is being hindered. Specific examples include the difficulty in distinguishing colleagues' faces due to fewer employees coming to the office and the need to wear masks, the increased opportunities for contact with new people due to the shift to a free address system, and the difficulty of responding to inquiries after work due to the increase in telecommuting. These factors have a negative impact on work efficiency and employee satisfaction (ES). A system that can efficiently resolve these issues is needed. [Means for solving the problem]

[0005] The present invention solves these problems with a system that includes a facial recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, and a data transmission and reception means. This system recognizes the face of the person in front of you and displays their profile information based on the data analyzed by the generative AI model. Voice recognition technology also makes it easy to search for employees and send messages. Furthermore, the AI ​​automatically responds after work is completed, reducing the workload and realizing smoother communication and improved work efficiency.

[0006] "Facial recognition means" is a function that analyzes image data acquired by a camera or sensor and identifies a person's face.

[0007] "Means for data analysis using generative AI models" refers to a function that analyzes data collected by AI models using machine learning and deep learning, and generates useful information.

[0008] "Operation means using voice recognition" is a function that analyzes the user's voice, converts it into text, and operates the system according to the instructions.

[0009] "Automatic response measures after work is completed" is a function that allows AI to automatically respond appropriately to inquiries even after the user has finished work.

[0010] "Data transmission / reception means" refers to a communication means for transmitting and receiving data between different devices and servers inside and outside the system. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0012] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0013] First, the terms used in the following description will be explained.

[0014] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0015] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0016] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0017] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0019] [First embodiment]

[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0021] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0024] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0027] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0031] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0032] The present invention, "Dokodemo Smile," is an information analysis tool that solves communication issues that arise due to changes in the working environment, and is composed of the following elements.

[0033] System Configuration

[0034] 1. Facial Recognition Methods

[0035] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[0036] 2. Data Analysis Methods Using Generative AI Models

[0037] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[0038] 3. Voice recognition operation

[0039] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[0040] 4. Automated response measures after business hours

[0041] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[0042] 5. Data transmission and reception means

[0043] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[0044] Program processing

[0045] The program processing of this system will be specifically explained below.

[0046] 1. Server Processing

[0047] Collection of user profiles: The server periodically collects data such as each user's name, department, work history, email content, etc. in cooperation with internal and external systems. This collected data is stored in a database.

[0048] Training the generative AI model: The server trains the generative AI model using the collected data and continues to retrain the model as new data is added.

[0049] Real-time data delivery: The server sends the latest profile information to the device in real time whenever a request is made.

[0050] 2. Terminal processing

[0051] Identifying the person in front of you: The device recognizes the face of the person in front of you from the camera image through the smart glasses and queries the server. It receives information from the server and displays the person's profile.

[0052] Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[0053] Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to generate optimal responses to inquiries and respond automatically.

[0054] 3. User Processing

[0055] Profile confirmation: Users can check the profile information of the person they are meeting through the screen of their smart glasses or device, enabling smooth communication.

[0056] Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, and sending messages.

[0057] Reduced workload: Users can reduce their workload by having AI respond to inquiries on their behalf after work or when they are off-duty.

[0058] Specific examples

[0059] For example, during a meeting, the smart glasses can recognize the face of the person they are meeting and display the person's name, department, and past email content. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[0060] summary

[0061] This invention, "Dokodemo Smile," is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automatic response after work, and data transmission and reception, in order to solve communication issues caused by changes in the work environment. This will realize smooth communication and improved work efficiency.

[0062] The processing flow will be explained below.

[0063] Step 1:

[0064] Collecting user profiles

[0065] The server connects with internal and external systems and periodically collects data such as each user's name, department, work history, email content, etc. This collected data is stored centrally in the server's database.

[0066] What happens: The server calls the API to collect the latest user information, converts it into an appropriate format, and stores it in the database.

[0067] Step 2:

[0068] Training generative AI models

[0069] The server preprocesses the collected data and trains generative AI models, and also retrains existing models as new data is added.

[0070] Specific operation: The server performs preprocessing such as data cleaning and tokenization, and then trains the AI ​​model using a GPU cluster.

[0071] Step 3:

[0072] Real-time data delivery

[0073] The server sends the latest profile information in real time in response to requests from the device.

[0074] Specific operation: The device sends a request to the server, and the server retrieves the user's information from the database and returns it to the device in JSON format.

[0075] Step 4:

[0076] Identifying the person you are facing

[0077] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, queries the server, receives information from the server, and displays the other person's profile.

[0078] What it does: The device runs a facial recognition algorithm, extracts facial features, and sends them to the server, which then returns the information.

[0079] Step 5:

[0080] Voice recognition operation

[0081] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages.

[0082] What it does: The device receives voice data, converts it to text using a speech recognition engine, and executes API requests based on that information.

[0083] Step 6:

[0084] Automating business operations with AI

[0085] The server uses the generation AI even after business hours or when the server is off-line to automatically generate optimal responses to inquiries.

[0086] How it works: The server passes the query to the AI ​​model, which generates an appropriate answer and replies via email or messaging system.

[0087] Step 7:

[0088] Profile confirmation

[0089] Users can check the profile information of the person they are meeting on the smart glasses display, enabling smooth communication.

[0090] Specific operation: The user decides what to talk about and what questions to ask based on the information displayed on the smart glasses.

[0091] Step 8:

[0092] Control with voice commands

[0093] Users can issue voice commands to the device to search for employees, make calls, send messages, and more.

[0094] Specific operation: The device recognizes the voice command given by the user and performs the appropriate operation.

[0095] Step 9:

[0096] Reduced workload

[0097] Users can reduce their workload by having AI take over after work or when they are off-duty.

[0098] Specific behavior: AI automatically responds to inquiries, reducing the tasks that users have to perform.

[0099] Example 1

[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0101] As the work environment changes, communication between employees is becoming less smooth. The constant wearing of masks and the introduction of a free address system can make facial recognition difficult, making face-to-face information sharing inconvenient. Furthermore, if work continues uninterrupted after the end of the working day, there is also the problem of infringing on employees' private time.

[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0103] In this invention, the server includes a face recognition means, a data analysis means using a machine learning model, an operation means for analyzing voice input, an automatic response means after work hours, and an information transmission and reception means. This allows employees to easily identify the person they are meeting and efficiently operate using voice commands. In addition, since the AI ​​automatically responds after work hours, employees can protect their private time.

[0104] "Facial identification means" refers to a technical means for identifying the face of the target person using video equipment.

[0105] "Means for data analysis using machine learning models" refers to technical means for performing data analysis using machine learning algorithms based on collected data.

[0106] "Operation means for analyzing voice input" refers to a technical means for analyzing a user's voice, converting it into text, and operating the system according to the commands.

[0107] "Automatic response measures after work is completed" refers to a technical measure that allows the generation AI to automatically respond appropriately to inquiries even after the user's work is completed.

[0108] "Means for transmitting and receiving information" refers to the technical means for transmitting and receiving information between different devices and servers inside and outside the system.

[0109] "Video equipment" is a general term for hardware devices used to capture images, such as cameras and sensors.

[0110] A "machine learning algorithm" is a mathematical model that learns patterns from large amounts of data and automates tasks such as prediction and classification.

[0111] A "voice command" is an instruction to operate the system using voice.

[0112] The present invention, "Dokodemo Smile," is an information analysis system that combines multiple technical means to solve communication issues that arise with changes in the work environment. The present invention is composed of the following main components:

[0113] Facial Identification Method

[0114] The device has the ability to identify the face of the person in front of it using video equipment such as cameras and sensors. Facial identification uses OpenCV and Microsoft® Azure® Face API. This method can eliminate the difficulty of face identification due to mask wearing and free address systems. For example, during a meeting, the camera in the smart glasses can detect the face of the person in front of it, convert the facial features into vectors, and send them to the server.

[0115] Data analysis methods using machine learning models

[0116] The server analyzes the collected data using a machine learning algorithm (for example, GPT-3 (registered trademark) or BERT). This data includes the user's name, department, work history, email content, etc., and is updated regularly. The analyzed information is sent to the device in real time whenever a request is made. For example, if a user wants to check the profile of a person they are meeting with, the server will provide the latest profile information.

[0117] A means of operation that analyzes voice input

[0118] The device analyzes the user's voice, converts it into text, and operates the system according to the commands. This voice recognition function uses Google® Cloud Speech-to-Text or Amazon Lex. Users can use voice commands to perform operations such as searching for employees, making calls, and sending messages. For example, if a user issues the voice command "View Tanaka's profile," the voice is converted into text and the device performs the corresponding operation.

[0119] Automated response measures after business hours

[0120] The device uses generative AI (e.g., GPT-3 or BERT) to generate optimal responses to inquiries after work hours or while in off-mode, and automatically responds. This allows users to spend their time with peace of mind after work. For example, if a user receives an inquiry email asking, "Please tell me more about tomorrow's meeting," while in off-mode, the AI ​​will automatically generate and send a reply.

[0121] Means of sending and receiving information

[0122] The server and terminals send and receive information between different devices and servers inside and outside the system. For this purpose, a real-time data streaming platform such as Apache Kafka is used. This makes it possible to obtain the necessary information in real time and display it on the terminal. For example, if new profile data is sent from the server during a meeting, it will be displayed on the terminal immediately.

[0123] Specific examples

[0124] For example, during a meeting, a user can recognize the face of the person they are meeting and display the person's name, department, and past email content on the smart glasses. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[0125] Examples of prompt statements

[0126] "Facial recognition is performed during meetings, and the other person's name, department, and past email content are displayed. The conversation proceeds smoothly, and documents can be emailed with subsequent voice commands. AI can automatically respond to inquiries even after work has finished."

[0127] In this way, the present invention, "Smile Anywhere," integrates a variety of technological means and responds to changes in the work environment, thereby realizing smooth communication and improved work efficiency.

[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0129] Program processing flow and specific explanation

[0130] Step 1: Collecting user profiles and storing them in a database

[0131] The server periodically collects data such as user name, department, work history, email content, etc. by linking with internal and external systems. This linking is done using REST API and SOAP.

[0132] Input: User name, department, work history, email content, etc. obtained through API calls

[0133] Data processing: Analyzes acquired data in JSON or XML format and converts it into a format that can be stored in a database.

[0134] Output: Correctly formatted user data stored in the database

[0135] What it does: Every night, the server uses a job scheduler (e.g., cron) to send a request to a specified API endpoint to retrieve the latest user data, which is then parsed and stored in a database.

[0136] Step 2: Train and update the generative AI model

[0137] The server uses the collected data to train a generative AI (e.g., GPT-3, BERT) and updates the model whenever new data is added.

[0138] Input: Latest user data stored in the database

[0139] Data computation: Using machine learning algorithms to train models and optimize parameters

[0140] Output: An updated generative AI model

[0141] How it works: Every weekend, the server begins retraining the model using newly collected data. Using a machine learning framework (e.g., TENSORFLOW® or PyTorch), the server learns patterns from large amounts of data and improves the model's accuracy.

[0142] Step 3: Identify your contact and view their profile

[0143] The device recognizes the face of the person it is meeting from the camera image via the smart glasses, queries the server and displays their profile.

[0144] Input: Smartglasses camera image

[0145] Data processing: Detecting face regions from video and generating face feature vectors

[0146] Output: Send the facial feature vector to the server, retrieve the corresponding person's profile information, and display it on the smart glasses.

[0147] How it works: During a meeting, the smart glasses analyze the camera footage in real time, use a facial recognition API to identify the face of the person they are meeting with, query the server to obtain the person's profile information, and display it on the smart glasses' display.

[0148] Step 4: Control with voice input

[0149] The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[0150] Input: User's voice

[0151] Data processing: Converting voice to text and parsing it as commands

[0152] Output: Based on the interpreted command, perform the corresponding operation on the terminal.

[0153] What happens: When a user says "View Tanaka's profile," the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the speech to text and take the appropriate action.

[0154] Step 5: Automated responses after the business day ends

[0155] The device responds to inquiries after work hours or in off-mode using the generative AI model, generating optimal answers and responding automatically.

[0156] Input: User inquiry (e.g., email)

[0157] Data Computation: Using generative AI models to generate relevant answers

[0158] Output: Send the generated answer to the user

[0159] Specific operation: When a user receives an inquiry email after work asking, "Please tell me the details about tomorrow's meeting," the device uses the generative AI model to automatically generate and send a response saying, "Tomorrow's meeting will be held in conference room A from 10:00."

[0160] Step 6: Send and receive information in real time

[0161] The server and terminal use a real-time data streaming platform such as Apache Kafka to send and receive information in real time.

[0162] Input: Newly acquired data and updated information

[0163] Data processing: process data in real time and send it to the device in the appropriate format

[0164] Output: Real-time updated information is displayed on the terminal.

[0165] Specific operation: When new profile data is sent from the server during a meeting, it is immediately displayed on the device, allowing the user to view the latest information.

[0166] In this way, the present invention "Smile Anywhere" smoothly solves communication issues that arise with changes in the work environment through multiple processing steps.

[0167] (Application example 1)

[0168] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0169] In recent years, there has been a demand for personalized service for each customer in brick-and-mortar stores. However, it is difficult to instantly grasp a customer's name, face, and purchase history, and there is a lack of systems to ensure smooth customer service. In addition, staff are required to respond to customer inquiries even after work hours, which increases the burden on staff. There is a need for an efficient customer service support system to solve these issues.

[0170] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0171] In this invention, the server includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after work is completed, a data transmission and reception unit, a unit for acquiring customer profile information, and a unit for analyzing purchase history. This enables personalized responses to each customer in a physical store, reducing the workload of staff and improving customer satisfaction.

[0172] "Facial recognition means" is a function that uses a camera or sensor to capture and identify a person's face.

[0173] "Data analysis means using generative AI models" is a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[0174] "Operation means using voice recognition" is a function that analyzes voice commands, converts them into text, and operates the system according to those commands.

[0175] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[0176] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[0177] "Means for obtaining customer profile information" refers to a function that collects data such as customer name, affiliation, purchase history, and survey responses, and stores it in a database.

[0178] "Means for analyzing purchasing history" refers to a function that analyzes a customer's past purchasing history and understands their preferences and trends.

[0179] The following detailed description of the embodiments of the present invention will be given. Note that the embodiments described herein embody the technical features included in the claims, but do not limit the technical scope of the invention.

[0180] A customer support system in a physical store mainly consists of three components: a server, a terminal, and a user. The roles and operations of each component are as follows:

[0181] 1. Server Roles and Operations

[0182] The server plays a central role in acquiring customer profile information and analyzing the data using generative AI models. The server operates using the following hardware and software:

[0183] Hardware: Server equipment with high-performance processors, memory, and large storage capacity.

[0184] Software: Database management systems (e.g., MySQL®), deep learning frameworks (e.g., TensorFlow), speech recognition APIs (e.g., Google Speech-to-Text).

[0185] The server performs the following process:

[0186] Collecting customer profiles: The server periodically collects data such as customer names, purchase history, and survey responses and stores it in a database.

[0187] Training the generative AI model: The server trains the generative AI model using the collected data and retrains the model whenever new data is added.

[0188] Real-time data delivery: The server sends the latest customer profile information to the device in real time upon request.

[0189] 2. Roles and Functions of the Device

[0190] Terminals are devices used by customer-facing staff, such as smart glasses and tablets, that operate using the following hardware and software:

[0191] Hardware: Smart glasses, camera, microphone.

[0192] Software: Facial recognition libraries (e.g., OpenCV), speech recognition software (e.g., Google Speech-to-Text).

[0193] The terminal performs the following process:

[0194] Customer face identification: The device recognizes the customer's face from the camera image of the smart glasses and queries the server for that information. The customer profile information sent from the server is displayed on the smart glasses display.

[0195] Voice recognition operation: The terminal recognizes voice commands and performs operations such as displaying customer information, checking inventory, and operating the cash register.

[0196] Automatic response after business hours: The device will use generative AI to automatically respond to inquiries after business hours or while in off mode.

[0197] 3. User Roles and Actions

[0198] The users of this system are the service staff who check customer information through the smart glasses or the terminal screen and perform operations using voice commands.

[0199] Check customer information: Staff can view information such as customer profiles and purchase history through the smart glasses screen.

[0200] Voice command operation: Staff use voice commands to search for information, check inventory, operate the cash register, and more.

[0201] Reduced workload: After the work is completed, AI automatically handles customer support, reducing the workload of staff.

[0202] Specific examples

[0203] For example, a sales staff member at a physical store can wear smart glasses and recognize the face of a customer who visits the store. The system displays the customer's profile and purchase history on the smart glasses' display, allowing the staff member to recommend products based on the customer's preferences and past purchase history. Furthermore, when the staff member issues a voice command such as "Show me recommended products for this customer," the system uses AI to select and display the appropriate products.

[0204] This will enable personalized service in physical stores, improving customer satisfaction. In addition, the AI ​​will automatically respond to customers after the store has finished its work, reducing the workload of staff.

[0205] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0206] Step 1:

[0207] The server collects customer profile information and stores it in a database. Specifically, it periodically obtains customer names, purchase histories, survey responses, etc. from the POS system in the physical store and the online survey system. The input to this process is data from the POS system and the survey system, and the output is customer information stored in the database.

[0208] Step 2:

[0209] The server uses the stored customer information to train a generative AI model. Specifically, it uses a deep learning framework (e.g., TensorFlow) to analyze customer purchasing trends and preferences. The input to this process is the customer information in the database, and the output is the trained AI model.

[0210] Step 3:

[0211] The server receives requests from the smart glasses and delivers real-time customer profile information. Specifically, it receives facial recognition data sent from the smart glasses' camera, retrieves the corresponding customer information from the database, and sends it to the smart glasses. The input of this process is facial recognition data, and the output is customer profile information.

[0212] Step 4:

[0213] The device (smart glasses) recognizes the customer's face using camera images. Specifically, it uses a facial recognition library (e.g., OpenCV) to extract facial features from the image and sends that information to a server. The input to this process is the camera image, and the output is facial recognition data.

[0214] Step 5:

[0215] The terminal recognizes the voice command and performs the required action. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text) to convert the voice command into text and follows the instructions to display customer information or check inventory. The input to this process is the voice command, and the output is the text command and the result of the action taken.

[0216] Step 6:

[0217] The terminal automatically responds to inquiries even after work hours. Specifically, it uses generative AI to generate appropriate responses to customer inquiries and automatically replies. The input to this process is the customer inquiry, and the output is the generated response.

[0218] Step 7:

[0219] The user checks customer information through the smart glasses screen and uses voice commands. Specifically, the user reads the information displayed on the smart glasses display and takes the necessary action to respond to the customer. The input of this process is the customer information displayed on the display, and the output is the user's specific action.

[0220] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0221] This invention is an information analysis tool that combines "Smile Anywhere" with an emotion engine to recognize user emotions and further improve smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[0222] System Configuration

[0223] 1. Facial Recognition Methods

[0224] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[0225] 2. Data Analysis Methods Using Generative AI Models

[0226] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[0227] 3. Voice recognition operation

[0228] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[0229] 4. Automated response measures after business hours

[0230] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[0231] 5. Data transmission and reception means

[0232] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[0233] 6. Emotion Engine

[0234] This function analyzes the user's voice and video data to recognize their emotional state, enabling them to respond and provide feedback according to their emotions.

[0235] Program processing

[0236] The system program of the present invention performs the following processes.

[0237] 1. Server Processing

[0238] - User profile collection: The server periodically collects data such as each user's name, department, work history, and email content in cooperation with internal and external systems. This data is then centralized in the server's database.

[0239] - Training generative AI models: The server trains generative AI models using collected data and continues to retrain existing models as new data is added.

[0240] - Training the emotion engine: The server trains the emotion engine based on the accumulated audio and video data.

[0241] - Real-time data delivery: The server sends the latest profile information and emotional state in real time in response to requests from the device.

[0242] 2. Terminal processing

[0243] - Face-to-face identification: The device recognizes the face of the person it is facing from the camera image through the smart glasses and queries the server. It receives information from the server and displays the face's profile and emotional state.

[0244] - Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees and sending messages. It also sends the voice data to an emotion engine for emotion recognition.

[0245] - Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to automatically generate the optimal response based on the inquiry and emotion.

[0246] 3. User Processing

[0247] - Profile Check: Users can check the profile information and emotional state of the person they are meeting through the smart glasses or device screen, enabling smooth communication.

[0248] - Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and also check the results of voice emotion analysis.

[0249] - Reduced workload: After work or while users are off-duty, AI can answer inquiries on their behalf and provide feedback based on emotion recognition, reducing their workload.

[0250] Specific examples

[0251] For example, during a meeting, a user can recognize the face of the person they are meeting and display the other person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, AI can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work.

[0252] summary

[0253] This invention is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automated responses after work is completed, data transmission and reception, and an emotion engine, to resolve communication barriers caused by changes in the work environment, thereby achieving smoother and more efficient communication and improved work efficiency.

[0254] The processing flow will be explained below.

[0255] Step 1:

[0256] Collecting user profiles

[0257] The server connects with internal and external systems and periodically collects each user's name, department, work history, email content, voice data, etc. This collected data is stored in a centralized database.

[0258] Specific operation: The server calls the API to collect the latest user information, converts it into an appropriate format, and saves it in the database.

[0259] Step 2:

[0260] Training generative AI models and emotion engines

[0261] The server preprocesses the collected data and trains the generative AI model and emotion engine, and retrains the AI ​​model and emotion engine every time new data is added.

[0262] How it works: The server performs preprocessing such as data cleaning and tokenization, and uses a GPU cluster to train the AI ​​model and emotion engine.

[0263] Step 3:

[0264] Real-time data delivery

[0265] The server transmits the latest profile information and emotional state in real time in response to requests from the device.

[0266] Specific operation: The device sends a request to the server, and the server retrieves the user's information and emotional state from the database and returns it to the device in JSON format.

[0267] Step 4:

[0268] Face-to-face identification and emotion recognition

[0269] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, and then uses an emotion engine to recognize the other person's emotional state. It then queries the server, receives the other person's profile information and emotional state, and displays them.

[0270] How it works: The device captures the face of the person it's facing with a camera, runs a facial recognition algorithm and emotion engine, and sends the extracted facial features and emotion information to the server, which then returns the relevant information.

[0271] Step 5:

[0272] Voice recognition operation

[0273] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages. The voice data is sent to an emotion engine to analyze the user's emotions.

[0274] What it does: The device receives voice data, converts it to text using a speech recognition engine, executes an API request based on that information, and analyzes the voice emotion using an emotion engine.

[0275] Step 6:

[0276] Automating business operations with AI

[0277] The server uses AI generation after work hours or when the user is off-duty to automatically generate the optimal response based on the inquiry and the user's emotional state.

[0278] How it works: The server passes the query and sentiment data to the model, generates the best answer, and sends it back via email or messaging system.

[0279] Step 7:

[0280] Profile and sentiment review

[0281] Users can check the profile information and emotional state of the person they are meeting through the screen of their smart glasses or device, allowing for smooth communication.

[0282] Specific behavior: The user sees the profile and emotional state of the other person displayed on the smart glasses and responds or engages in conversation appropriately.

[0283] Step 8:

[0284] Control and confirm emotions with voice commands

[0285] Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and checking their own and the other person's emotional state.

[0286] Specific operation: The user issues a voice command, the device recognizes it, and performs the necessary operation, displaying emotional information analyzed by the emotion engine.

[0287] Step 9:

[0288] Reduced workload

[0289] The AI ​​can take over after work or while the user is off-duty, reducing the burden on the user. Feedback based on emotion recognition also reduces stress and enables appropriate work responses.

[0290] Specific operation: AI automatically generates and sends appropriate responses to inquiries. It analyzes emotional information and provides appropriate feedback to users.

[0291] Example 2

[0292] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0293] In recent years, changes in the work environment have made it increasingly difficult for employees to communicate smoothly. Furthermore, remote work and flexible office settings have reduced face-to-face communication, making it difficult to recognize emotions and streamline work. This can lead to delays in work progress and team cooperation. A system is needed to resolve these issues, promote smooth communication between employees, and improve work efficiency.

[0294] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0295] In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, an emotion engine means, a user profile collection means, a generative AI model training means, an emotion engine training means, a real-time data distribution means, a face-to-face person identification means, and an AI-based automation means for work responses. This enables smooth communication between employees and improves work efficiency.

[0296] "Facial recognition means" is a technology that uses cameras and sensors to identify a person's face.

[0297] "Data analysis methods using generative AI models" is a technology that utilizes AI models that use machine learning and deep learning to analyze collected data and generate useful information.

[0298] "Operation means using voice recognition" is a technology that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[0299] "Automatic response measures after work is completed" refers to technology that enables the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[0300] "Data transmission / reception means" refers to technology for transmitting and receiving data between different devices and servers inside and outside the system.

[0301] The "emotion engine means" is a technology for analyzing the user's voice and video data and recognizing the user's emotional state.

[0302] "Means for collecting user profiles" refers to technology that allows the server to collect data such as the user's name, department, work history, and email content from systems both inside and outside the company.

[0303] A "generative AI model training means" is a technique for training a generative AI model using collected data and for retraining the model whenever new data is added.

[0304] "Means for training the emotion engine" refers to a technology for training the emotion engine using stored audio and video data.

[0305] "Real-time data distribution means" refers to technology for transmitting the latest user profile information and emotional state in real time in response to a request from a terminal.

[0306] The "means for identifying the person being met" is a technology that recognizes the face of the person being met from camera images via smart glasses and queries the server.

[0307] "AI-based automation of business operations" is a technology that uses generative AI after work hours or during off-duty periods to automatically generate optimal responses based on inquiries and emotions.

[0308] This invention is an information analysis tool that recognizes user emotions by combining an emotion engine with "Smile Anywhere," thereby improving smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[0309] System Configuration

[0310] Facial Recognition Methods

[0311] The system uses cameras and sensors to capture the face of the person facing you and identify them. This method makes it possible to identify the other person even in situations where facial identification is difficult due to mask wearing or free address systems. Specifically, it utilizes face recognition technology (e.g., OpenCV) using camera images.

[0312] Data analysis methods using generative AI models

[0313] The system uses AI models (e.g., GPT-4 (registered trademark)) that use machine learning and deep learning to analyze the collected data and generate useful information. This method allows it to present relevant information such as the other party's work history and email content.

[0314] Voice recognition operation

[0315] The system analyzes the user's voice, converts it into text, and operates the system according to the commands. This method allows hands-free employee search and message sending. Specifically, it uses a voice recognition service (e.g., Google Assistant).

[0316] Automated response measures after business hours

[0317] The system uses generative AI to automatically respond to inquiries after users have finished their work, giving users more freedom to use their time after work.

[0318] Data transmission and reception means

[0319] Send and receive data between different devices and servers inside and outside the system. This method makes it possible to obtain the necessary information in real time and display it on the appropriate device. Specifically, it uses RESTful APIs.

[0320] Emotion Engine

[0321] The system analyzes the user's voice and video data to recognize their emotional state, enabling it to respond and provide feedback according to their emotions.

[0322] How user profiles are collected

[0323] The server connects with internal and external systems to periodically collect data such as each user's name, department, work history, and email content, and centralizes it in a database. Specifically, it obtains data from LDAP directories and email servers via API.

[0324] A means of training generative AI models

[0325] The server uses the collected data to train a generative AI model, and continues to retrain the model as new data is added.

[0326] A means of training the emotion engine

[0327] The server trains the emotion engine based on the stored audio and video data. Specifically, it converts the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeds it into the emotion analysis model.

[0328] Real-time data delivery method

[0329] The server receives requests from the device and sends the latest user profile information and emotional state in real time. Specifically, it receives requests from the device using a RESTful API and returns them in JSON format.

[0330] Means of identifying the person you are facing

[0331] The device recognizes the face of the person in front of it from the camera image through the smart glasses, queries the server, and receives information from the server to display the other person's profile and emotional state. Specifically, it uses facial recognition software (e.g., OpenCV).

[0332] Automating business processes using AI

[0333] The device uses generative AI to automatically generate the optimal response based on the inquiry and the customer's emotion after work hours or when the device is in off mode. The AI ​​model is called via API, and the inquiry content and emotional data are input, and the generated response is automatically sent via email or chat.

[0334] Specific examples

[0335] For example, a user can recognize the face of the person they are meeting with during a meeting and display the person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, the AI ​​can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work. An example prompt might be, "Please describe a system that allows a user to identify the face of the person they are meeting with during a meeting and display the other person's emotional state and work history."

[0336] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0337] Step 1:

[0338] Collecting user profiles

[0339] The server connects with internal and external systems (for example, LDAP directories and mail servers) to collect data such as each user's name, department, work history, email content, etc. This data is obtained via API and centralized in the server's database.

[0340] Input: Data from an LDAP directory or mail server.

[0341] Output: Centralized database.

[0342] Step 2:

[0343] Training generative AI models

[0344] The server uses the collected data to train a generative AI model (e.g., GPT-4), and retrains this model to keep it up to date whenever new data is added.

[0345] Input: Information from a centralized database.

[0346] Output: A trained generative AI model.

[0347] Step 3:

[0348] Training the Emotional Engine

[0349] The server trains the emotion engine based on the stored audio and video data, for example by converting the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeding the text data into the emotion analysis model.

[0350] Input: Audio data, video data.

[0351] Output: A trained emotion engine.

[0352] Step 4:

[0353] Identifying the person you are facing

[0354] The device recognizes the face of the person it is meeting from the camera image through the smart glasses. Using facial recognition software (e.g., OpenCV), it sends the recognized facial data to the server, which then returns the other person's profile and emotional state.

[0355] Input: Camera footage.

[0356] Output: The other person's profile information and emotional state.

[0357] Step 5:

[0358] Viewing a User Profile

[0359] Users can check the profile information and emotional state of the person they are meeting through the smart glasses, which allows for smooth communication. The information is displayed on the smart glasses' HUD.

[0360] Input: Profile information and emotional state received from the server.

[0361] Output: Information displayed on the smart glasses HUD.

[0362] Step 6:

[0363] Voice recognition operation

[0364] The device uses a voice recognition function to analyze the user's voice commands. For example, it uses a voice recognition service (e.g., Google Assistant) to convert the voice commands into text and analyzes the text data. As a result, it becomes possible to perform operations such as searching for employees and sending messages.

[0365] Input: The user's voice command.

[0366] Output: Actions based on the parsed voice command.

[0367] Step 7:

[0368] Real-time data delivery

[0369] The server receives requests from the device and sends the latest user profile information and emotional state in real time. It uses a RESTful API to receive requests, retrieve the latest information from the database, and return it in JSON format.

[0370] Input: A request from the terminal.

[0371] Output: Latest user profile information and emotional state.

[0372] Step 8:

[0373] Automating business operations with AI

[0374] The device uses generative AI after work hours or when in off-duty mode to automatically generate the optimal response based on the inquiry and emotion. For example, the AI ​​model can be called via API, the inquiry content and emotion data are input, and the generated response is automatically sent via email or chat.

[0375] Input: Enquiry content and emotion data.

[0376] Output: The generated answer.

[0377] Step 9:

[0378] Reduced workload

[0379] The AI ​​will answer inquiries on behalf of the user after work or while the user is off-duty. It also reduces the burden of work by providing feedback based on emotion recognition. For example, the user can request, "Please send me an email with your thoughts on today's meeting," and the AI ​​will automatically respond.

[0380] Input: User request.

[0381] Output: Auto-generated emails and feedback.

[0382] (Application example 2)

[0383] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0384] Customer service in modern brick-and-mortar stores requires quick and effective responses to individual customers, but achieving this is not easy. For example, it is difficult for store employees to check inventory while serving a customer, or to appropriately understand and respond to a customer's emotional state. Furthermore, employees are required to respond quickly to customer inquiries even after their shifts have finished, placing a heavy burden on them. These challenges hinder efficient customer service and the improvement of business processes.

[0385] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, a customer information display means using emotion recognition, and an inventory confirmation means using voice operation. This allows store employees to not only instantly check customer information and emotional state when meeting a customer, but also to quickly check inventory and suggest products using voice commands. Furthermore, realizing automatic response after work is completed using AI reduces the burden on employees and enables more efficient customer service.

[0386] "Facial recognition means" is a function that uses a camera or sensor to capture a person's face and identify a specific individual.

[0387] "Means for data analysis using generative AI models" refers to a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[0388] "Operation means using voice recognition" is a function that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[0389] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[0390] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[0391] The "means for displaying customer information based on emotion recognition" is a function that analyzes the customer's voice and video data, recognizes their emotional state, and displays that information.

[0392] The "voice-operated inventory check means" is a function that allows you to check the inventory status using voice commands.

[0393] System Overview

[0394] The present invention is a system for improving the efficiency of customer service in brick-and-mortar stores. This system includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after business hours are completed, a data transmission and reception unit, a customer information display unit using emotion recognition, and an inventory confirmation unit using voice operation.

[0395] Hardware and Software Use

[0396] The system uses the following hardware and software:

[0397] Hardware

[0398] Smart glasses (e.g., Google Glass(R), Vuzix Blade): Equipped with a camera and microphone to capture the face of the person facing you.

[0399] High-performance camera: Captures video data for facial recognition.

[0400] Microphone: Captures voice commands and is used for voice recognition functions.

[0401] software

[0402] Operating system: ANDROID (registered trademark), iOS

[0403] Generative AI model: OpenAI (registered trademark), data analysis using GPT-4

[0404] Emotion recognition engine: Affectiva

[0405] Speech Recognition API: Google Speech-to-Text

[0406] Cloud database: Firebase, AWS (registered trademark) RDS

[0407] Real-time data transfer: WebSocket, Firebase Realtime Database

[0408] Program processing

[0409] Acquiring customer information through facial recognition

[0410] The camera in the smart glasses captures the customer's face. The captured image data is analyzed by facial recognition and compared with a database to obtain the customer's profile information. This allows the necessary information to be displayed immediately when interacting with the customer.

[0411] Emotion-aware response

[0412] The captured video data is sent to an emotion recognition engine (Affectiva) to analyze the customer's emotional state. The resulting emotional state (e.g., satisfaction, confusion, interest, etc.) is displayed on the smart glasses' display, allowing store staff to quickly respond according to the customer's emotions.

[0413] Voice recognition operation

[0414] When an employee issues a voice command, the system uses the Google Speech-to-Text API to convert the voice into text, analyzes it, and performs the command accordingly. For example, if you issue the voice command "Check inventory, product number 12345," inventory information will be retrieved from a cloud database and displayed in real time.

[0415] Automatic response after business hours

[0416] Even after work hours, automated responses are provided using generative AI models (OpenAI, GPT-4). The AI ​​model can generate and send appropriate responses to customer inquiries, reducing the workload of employees.

[0417] Examples and prompts

[0418] Example 1: Checking inventory

[0419] When a store clerk issues a voice command through the smart glasses, such as "Check stock, product number 12345," the system retrieves the product's stock status in real time from the cloud database and displays it on the smart glasses' display.

[0420] Example 2: Customer service using sentiment analysis

[0421] An emotion recognition engine analyzes the video data captured by the smart glasses' camera, and if it indicates that the customer is confused, employees can immediately change their response.

[0422] ChatGPT(R) prompt example

[0423] Prompt for generative AI models (OpenAI, GPT-4):

[0424] "The customer is confused. What should we do next?"

[0425] Customer profile information:

[0426] Name: Taro

[0427] Past purchase history: Many home appliances

[0428] Preferences: Products with the latest technology

[0429] Emotional state: Confused

[0430] Generation example:

[0431] "Taro, what kind of home appliances would you like to buy today? We also have the latest smart home devices."

[0432] In this way, the operation of the system can significantly improve the quality of customer service and business efficiency.

[0433] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0434] Step 1:

[0435] A user puts on the smart glasses and launches the application. The smart glasses' camera captures the customer's face, and the video data is input into the facial recognition means. The facial recognition model analyzes the video data and extracts facial feature points. The extracted feature points are sent to the server for matching with a database. The server searches the database, obtains the corresponding customer's profile information, and sends it to the terminal. The terminal displays this information on the smart glasses' display. This allows employees to check information such as the customer's name and purchase history in real time.

[0436] Step 2:

[0437] The customer's video data captured by the smart glasses is input into an emotion recognition engine (Affectiva). The emotion recognition engine analyzes the customer's facial expressions and identifies their emotional state. The identified emotional information (e.g., satisfaction, confusion, interest, etc.) is sent to the terminal and displayed on the smart glasses' display. This allows employees to quickly respond according to the customer's emotions.

[0438] Step 3:

[0439] When a user issues a voice command, the microphone in the smart glasses captures the voice. The captured voice data is input into a speech recognition API (Google Speech-to-Text) and converted into text. The converted text data is sent to the server, where it is analyzed by a generative AI model (OpenAI, GPT-4). For example, in response to the command "Check stock, product number 12345," the server searches a cloud database and retrieves stock information for product number 12345. This information is sent to the device and displayed on the smart glasses' display. This allows employees to respond quickly to customer questions without taking their eyes off the customer.

[0440] Step 4:

[0441] Even after the user has finished their work, the generative AI model (OpenAI, GPT-4) remains in standby mode on the server. When a customer makes an inquiry, the inquiry is entered into the server. The server uses the generative AI model to generate the optimal answer and sends it to the customer via an automated response system. This enables prompt customer responses even after hours, reducing the workload of employees.

[0442] Through these steps, the present invention is a system that can significantly improve the efficiency of customer service in physical stores and improve business processes.

[0443] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0444] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0445] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0446] [Second embodiment]

[0447] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0448] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0449] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0450] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0451] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0453] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0454] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0455] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0456] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0457] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0458] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0459] The present invention, "Dokodemo Smile," is an information analysis tool that solves communication issues that arise due to changes in the working environment, and is composed of the following elements.

[0460] System Configuration

[0461] 1. Facial Recognition Methods

[0462] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[0463] 2. Data Analysis Methods Using Generative AI Models

[0464] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[0465] 3. Voice recognition operation

[0466] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[0467] 4. Automated response measures after business hours

[0468] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[0469] 5. Data transmission and reception means

[0470] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[0471] Program processing

[0472] The program processing of this system will be specifically explained below.

[0473] 1. Server Processing

[0474] Collection of user profiles: The server periodically collects data such as each user's name, department, work history, email content, etc. in cooperation with internal and external systems. This collected data is stored in a database.

[0475] Training the generative AI model: The server trains the generative AI model using the collected data and continues to retrain the model as new data is added.

[0476] Real-time data delivery: The server sends the latest profile information to the device in real time whenever a request is made.

[0477] 2. Terminal processing

[0478] Identifying the person in front of you: The device recognizes the face of the person in front of you from the camera image through the smart glasses and queries the server. It receives information from the server and displays the person's profile.

[0479] Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[0480] Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to generate optimal responses to inquiries and respond automatically.

[0481] 3. User Processing

[0482] Profile confirmation: Users can check the profile information of the person they are meeting through the screen of their smart glasses or device, enabling smooth communication.

[0483] Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, and sending messages.

[0484] Reduced workload: Users can reduce their workload by having AI respond to inquiries on their behalf after work or when they are off-duty.

[0485] Specific examples

[0486] For example, during a meeting, the smart glasses can recognize the face of the person they are meeting and display the person's name, department, and past email content. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[0487] summary

[0488] This invention, "Dokodemo Smile," is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automatic response after work, and data transmission and reception, in order to solve communication issues caused by changes in the work environment. This will realize smooth communication and improved work efficiency.

[0489] The processing flow will be explained below.

[0490] Step 1:

[0491] Collecting user profiles

[0492] The server connects with internal and external systems and periodically collects data such as each user's name, department, work history, email content, etc. This collected data is stored centrally in the server's database.

[0493] What happens: The server calls the API to collect the latest user information, converts it into an appropriate format, and stores it in the database.

[0494] Step 2:

[0495] Training generative AI models

[0496] The server preprocesses the collected data and trains generative AI models, and also retrains existing models as new data is added.

[0497] Specific operation: The server performs preprocessing such as data cleaning and tokenization, and then trains the AI ​​model using a GPU cluster.

[0498] Step 3:

[0499] Real-time data delivery

[0500] The server sends the latest profile information in real time in response to requests from the device.

[0501] Specific operation: The device sends a request to the server, and the server retrieves the user's information from the database and returns it to the device in JSON format.

[0502] Step 4:

[0503] Identifying the person you are facing

[0504] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, queries the server, receives information from the server, and displays the other person's profile.

[0505] What it does: The device runs a facial recognition algorithm, extracts facial features, and sends them to the server, which then returns the information.

[0506] Step 5:

[0507] Voice recognition operation

[0508] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages.

[0509] What it does: The device receives voice data, converts it to text using a speech recognition engine, and executes API requests based on that information.

[0510] Step 6:

[0511] Automating business operations with AI

[0512] The server uses the generation AI even after business hours or when the server is off-line to automatically generate optimal responses to inquiries.

[0513] How it works: The server passes the query to the AI ​​model, which generates an appropriate answer and replies via email or messaging system.

[0514] Step 7:

[0515] Profile confirmation

[0516] Users can check the profile information of the person they are meeting on the smart glasses display, enabling smooth communication.

[0517] Specific operation: The user decides what to talk about and what questions to ask based on the information displayed on the smart glasses.

[0518] Step 8:

[0519] Control with voice commands

[0520] Users can issue voice commands to the device to search for employees, make calls, send messages, and more.

[0521] Specific operation: The device recognizes the voice command given by the user and performs the appropriate operation.

[0522] Step 9:

[0523] Reduced workload

[0524] Users can reduce their workload by having AI take over after work or when they are off-duty.

[0525] Specific behavior: AI automatically responds to inquiries, reducing the tasks that users have to perform.

[0526] Example 1

[0527] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0528] As the work environment changes, communication between employees is becoming less smooth. The constant wearing of masks and the introduction of a free address system can make facial recognition difficult, making face-to-face information sharing inconvenient. Furthermore, if work continues uninterrupted after the end of the working day, there is also the problem of infringing on employees' private time.

[0529] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0530] In this invention, the server includes a face recognition means, a data analysis means using a machine learning model, an operation means for analyzing voice input, an automatic response means after work hours, and an information transmission and reception means. This allows employees to easily identify the person they are meeting and efficiently operate using voice commands. In addition, since the AI ​​automatically responds after work hours, employees can protect their private time.

[0531] "Facial identification means" refers to a technical means for identifying the face of the target person using video equipment.

[0532] "Means for data analysis using machine learning models" refers to technical means for performing data analysis using machine learning algorithms based on collected data.

[0533] "Operation means for analyzing voice input" refers to a technical means for analyzing a user's voice, converting it into text, and operating the system according to the commands.

[0534] "Automatic response measures after work is completed" refers to a technical measure that allows the generation AI to automatically respond appropriately to inquiries even after the user's work is completed.

[0535] "Means for transmitting and receiving information" refers to the technical means for transmitting and receiving information between different devices and servers inside and outside the system.

[0536] "Video equipment" is a general term for hardware devices used to capture images, such as cameras and sensors.

[0537] A "machine learning algorithm" is a mathematical model that learns patterns from large amounts of data and automates tasks such as prediction and classification.

[0538] A "voice command" is an instruction to operate the system using voice.

[0539] The present invention, "Dokodemo Smile," is an information analysis system that combines multiple technical means to solve communication issues that arise with changes in the work environment. The present invention is composed of the following main components:

[0540] Facial Identification Method

[0541] The device has the ability to identify the face of the person in front of it using video equipment such as cameras and sensors. Facial identification uses OpenCV and Microsoft Azure's Face API. This method can eliminate the difficulty of face identification due to mask wearing and free address systems. For example, during a meeting, the camera in the smart glasses can detect the face of the person in front of it, convert the facial features into vectors, and send them to a server.

[0542] Data analysis methods using machine learning models

[0543] The server analyzes the collected data using machine learning algorithms (e.g., GPT-3 or BERT). This data includes the user's name, department, work history, email content, etc., and is updated regularly. The analyzed information is sent to the device in real time whenever a request is made. For example, if a user wants to check the profile of a person they are meeting with, the server will provide the latest profile information.

[0544] A means of operation that analyzes voice input

[0545] The device analyzes the user's voice, converts it into text, and operates the system according to the commands. This voice recognition function uses Google Cloud Speech-to-Text and Amazon Lex. Users can use voice commands to perform operations such as searching for employees, making calls, and sending messages. For example, if you issue the voice command "View Tanaka's profile," the voice will be converted into text and the device will perform the corresponding operation.

[0546] Automated response measures after business hours

[0547] The device uses generative AI (e.g., GPT-3 or BERT) to generate optimal responses to inquiries after work hours or while in off-mode, and automatically responds. This allows users to spend their time with peace of mind after work. For example, if a user receives an inquiry email asking, "Please tell me more about tomorrow's meeting," while in off-mode, the AI ​​will automatically generate and send a reply.

[0548] Means of sending and receiving information

[0549] The server and terminals send and receive information between different devices and servers inside and outside the system. For this purpose, a real-time data streaming platform such as Apache Kafka is used. This makes it possible to obtain the necessary information in real time and display it on the terminal. For example, if new profile data is sent from the server during a meeting, it will be displayed on the terminal immediately.

[0550] Specific examples

[0551] For example, during a meeting, a user can recognize the face of the person they are meeting and display the person's name, department, and past email content on the smart glasses. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[0552] Examples of prompt statements

[0553] "Facial recognition is performed during meetings, and the other person's name, department, and past email content are displayed. The conversation proceeds smoothly, and documents can be emailed with subsequent voice commands. AI can automatically respond to inquiries even after work has finished."

[0554] In this way, the present invention, "Smile Anywhere," integrates a variety of technological means and responds to changes in the work environment, thereby realizing smooth communication and improved work efficiency.

[0555] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0556] Program processing flow and specific explanation

[0557] Step 1: Collecting user profiles and storing them in a database

[0558] The server periodically collects data such as user name, department, work history, email content, etc. by linking with internal and external systems. This linking is done using REST API and SOAP.

[0559] Input: User name, department, work history, email content, etc. obtained through API calls

[0560] Data processing: Analyzes acquired data in JSON or XML format and converts it into a format that can be stored in a database.

[0561] Output: Correctly formatted user data stored in the database

[0562] What it does: Every night, the server uses a job scheduler (e.g., cron) to send a request to a specified API endpoint to retrieve the latest user data, which is then parsed and stored in a database.

[0563] Step 2: Train and update the generative AI model

[0564] The server uses the collected data to train a generative AI (e.g., GPT-3, BERT) and updates the model whenever new data is added.

[0565] Input: Latest user data stored in the database

[0566] Data computation: Using machine learning algorithms to train models and optimize parameters

[0567] Output: An updated generative AI model

[0568] How it works: Every weekend, the server begins retraining the model using newly collected data. It uses a machine learning framework (e.g., TensorFlow or PyTorch) to learn patterns from large amounts of data and improve the model's accuracy.

[0569] Step 3: Identify your contact and view their profile

[0570] The device recognizes the face of the person it is meeting from the camera image via the smart glasses, queries the server and displays their profile.

[0571] Input: Smartglasses camera image

[0572] Data processing: Detecting face regions from video and generating face feature vectors

[0573] Output: Send the facial feature vector to the server, retrieve the corresponding person's profile information, and display it on the smart glasses.

[0574] How it works: During a meeting, the smart glasses analyze the camera footage in real time, use a facial recognition API to identify the face of the person they are meeting with, query the server to obtain the person's profile information, and display it on the smart glasses' display.

[0575] Step 4: Control with voice input

[0576] The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[0577] Input: User's voice

[0578] Data processing: Converting voice to text and parsing it as commands

[0579] Output: Based on the interpreted command, perform the corresponding operation on the terminal.

[0580] What happens: When a user says "View Tanaka's profile," the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the speech to text and take the appropriate action.

[0581] Step 5: Automated responses after the business day ends

[0582] The device responds to inquiries after work hours or in off-mode using the generative AI model, generating optimal answers and responding automatically.

[0583] Input: User inquiry (e.g., email)

[0584] Data Computation: Using generative AI models to generate relevant answers

[0585] Output: Send the generated answer to the user

[0586] Specific operation: When a user receives an inquiry email after work asking, "Please tell me the details about tomorrow's meeting," the device uses the generative AI model to automatically generate and send a response saying, "Tomorrow's meeting will be held in conference room A from 10:00."

[0587] Step 6: Send and receive information in real time

[0588] The server and terminal use a real-time data streaming platform such as Apache Kafka to send and receive information in real time.

[0589] Input: Newly acquired data and updated information

[0590] Data processing: process data in real time and send it to the device in the appropriate format

[0591] Output: Real-time updated information is displayed on the terminal.

[0592] Specific operation: When new profile data is sent from the server during a meeting, it is immediately displayed on the device, allowing the user to view the latest information.

[0593] In this way, the present invention "Smile Anywhere" smoothly solves communication issues that arise with changes in the work environment through multiple processing steps.

[0594] (Application example 1)

[0595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0596] In recent years, there has been a demand for personalized service for each customer in brick-and-mortar stores. However, it is difficult to instantly grasp a customer's name, face, and purchase history, and there is a lack of systems to ensure smooth customer service. In addition, staff are required to respond to customer inquiries even after work hours, which increases the burden on staff. There is a need for an efficient customer service support system to solve these issues.

[0597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0598] In this invention, the server includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after work is completed, a data transmission and reception unit, a unit for acquiring customer profile information, and a unit for analyzing purchase history. This enables personalized responses to each customer in a physical store, reducing the workload of staff and improving customer satisfaction.

[0599] "Facial recognition means" is a function that uses a camera or sensor to capture and identify a person's face.

[0600] "Data analysis means using generative AI models" is a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[0601] "Operation means using voice recognition" is a function that analyzes voice commands, converts them into text, and operates the system according to those commands.

[0602] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[0603] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[0604] "Means for obtaining customer profile information" refers to a function that collects data such as customer name, affiliation, purchase history, and survey responses, and stores it in a database.

[0605] "Means for analyzing purchasing history" refers to a function that analyzes a customer's past purchasing history and understands their preferences and trends.

[0606] The following detailed description of the embodiments of the present invention will be given. Note that the embodiments described herein embody the technical features included in the claims, but do not limit the technical scope of the invention.

[0607] A customer support system in a physical store mainly consists of three components: a server, a terminal, and a user. The roles and operations of each component are as follows:

[0608] 1. Server Roles and Operations

[0609] The server plays a central role in acquiring customer profile information and analyzing the data using generative AI models. The server operates using the following hardware and software:

[0610] Hardware: Server equipment with high-performance processors, memory, and large storage capacity.

[0611] Software: Database management systems (e.g., MySQL), deep learning frameworks (e.g., TensorFlow), speech recognition APIs (e.g., Google Speech-to-Text).

[0612] The server performs the following process:

[0613] Collecting customer profiles: The server periodically collects data such as customer names, purchase history, and survey responses and stores it in a database.

[0614] Training the generative AI model: The server trains the generative AI model using the collected data and retrains the model whenever new data is added.

[0615] Real-time data delivery: The server sends the latest customer profile information to the device in real time upon request.

[0616] 2. Roles and Functions of the Device

[0617] Terminals are devices used by customer-facing staff, such as smart glasses and tablets, that operate using the following hardware and software:

[0618] Hardware: Smart glasses, camera, microphone.

[0619] Software: Facial recognition libraries (e.g., OpenCV), speech recognition software (e.g., Google Speech-to-Text).

[0620] The terminal performs the following process:

[0621] Customer face identification: The device recognizes the customer's face from the camera image of the smart glasses and queries the server for that information. The customer profile information sent from the server is displayed on the smart glasses display.

[0622] Voice recognition operation: The terminal recognizes voice commands and performs operations such as displaying customer information, checking inventory, and operating the cash register.

[0623] Automatic response after business hours: The device will use generative AI to automatically respond to inquiries after business hours or while in off mode.

[0624] 3. User Roles and Actions

[0625] The users of this system are the service staff who check customer information through the smart glasses or the terminal screen and perform operations using voice commands.

[0626] Check customer information: Staff can view information such as customer profiles and purchase history through the smart glasses screen.

[0627] Voice command operation: Staff use voice commands to search for information, check inventory, operate the cash register, and more.

[0628] Reduced workload: After the work is completed, AI automatically handles customer support, reducing the workload of staff.

[0629] Specific examples

[0630] For example, a sales staff member at a physical store can wear smart glasses and recognize the face of a customer who visits the store. The system displays the customer's profile and purchase history on the smart glasses' display, allowing the staff member to recommend products based on the customer's preferences and past purchase history. Furthermore, when the staff member issues a voice command such as "Show me recommended products for this customer," the system uses AI to select and display the appropriate products.

[0631] This will enable personalized service in physical stores, improving customer satisfaction. In addition, the AI ​​will automatically respond to customers after the store has finished its work, reducing the workload of staff.

[0632] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0633] Step 1:

[0634] The server collects customer profile information and stores it in a database. Specifically, it periodically obtains customer names, purchase histories, survey responses, etc. from the POS system in the physical store and the online survey system. The input to this process is data from the POS system and the survey system, and the output is customer information stored in the database.

[0635] Step 2:

[0636] The server uses the stored customer information to train a generative AI model. Specifically, it uses a deep learning framework (e.g., TensorFlow) to analyze customer purchasing trends and preferences. The input to this process is the customer information in the database, and the output is the trained AI model.

[0637] Step 3:

[0638] The server receives requests from the smart glasses and delivers real-time customer profile information. Specifically, it receives facial recognition data sent from the smart glasses' camera, retrieves the corresponding customer information from the database, and sends it to the smart glasses. The input of this process is facial recognition data, and the output is customer profile information.

[0639] Step 4:

[0640] The device (smart glasses) recognizes the customer's face using camera images. Specifically, it uses a facial recognition library (e.g., OpenCV) to extract facial features from the image and sends that information to a server. The input to this process is the camera image, and the output is facial recognition data.

[0641] Step 5:

[0642] The terminal recognizes the voice command and performs the required action. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text) to convert the voice command into text and follows the instructions to display customer information or check inventory. The input to this process is the voice command, and the output is the text command and the result of the action taken.

[0643] Step 6:

[0644] The terminal automatically responds to inquiries even after work hours. Specifically, it uses generative AI to generate appropriate responses to customer inquiries and automatically replies. The input to this process is the customer inquiry, and the output is the generated response.

[0645] Step 7:

[0646] The user checks customer information through the smart glasses screen and uses voice commands. Specifically, the user reads the information displayed on the smart glasses display and takes the necessary action to respond to the customer. The input of this process is the customer information displayed on the display, and the output is the user's specific action.

[0647] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0648] This invention is an information analysis tool that combines "Smile Anywhere" with an emotion engine to recognize user emotions and further improve smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[0649] System Configuration

[0650] 1. Facial Recognition Methods

[0651] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[0652] 2. Data Analysis Methods Using Generative AI Models

[0653] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[0654] 3. Voice recognition operation

[0655] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[0656] 4. Automated response measures after business hours

[0657] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[0658] 5. Data transmission and reception means

[0659] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[0660] 6. Emotion Engine

[0661] This function analyzes the user's voice and video data to recognize their emotional state, enabling them to respond and provide feedback according to their emotions.

[0662] Program processing

[0663] The system program of the present invention performs the following processes.

[0664] 1. Server Processing

[0665] - User profile collection: The server periodically collects data such as each user's name, department, work history, and email content in cooperation with internal and external systems. This data is then centralized in the server's database.

[0666] - Training generative AI models: The server trains generative AI models using collected data and continues to retrain existing models as new data is added.

[0667] - Training the emotion engine: The server trains the emotion engine based on the accumulated audio and video data.

[0668] - Real-time data delivery: The server sends the latest profile information and emotional state in real time in response to requests from the device.

[0669] 2. Terminal processing

[0670] - Face-to-face identification: The device recognizes the face of the person it is facing from the camera image through the smart glasses and queries the server. It receives information from the server and displays the face's profile and emotional state.

[0671] - Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees and sending messages. It also sends the voice data to an emotion engine for emotion recognition.

[0672] - Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to automatically generate the optimal response based on the inquiry and emotion.

[0673] 3. User Processing

[0674] - Profile Check: Users can check the profile information and emotional state of the person they are meeting through the smart glasses or device screen, enabling smooth communication.

[0675] - Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and also check the results of voice emotion analysis.

[0676] - Reduced workload: After work or while users are off-duty, AI can answer inquiries on their behalf and provide feedback based on emotion recognition, reducing their workload.

[0677] Specific examples

[0678] For example, during a meeting, a user can recognize the face of the person they are meeting and display the other person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, AI can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work.

[0679] summary

[0680] This invention is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automated responses after work is completed, data transmission and reception, and an emotion engine, to resolve communication barriers caused by changes in the work environment, thereby achieving smoother and more efficient communication and improved work efficiency.

[0681] The processing flow will be explained below.

[0682] Step 1:

[0683] Collecting user profiles

[0684] The server connects with internal and external systems and periodically collects each user's name, department, work history, email content, voice data, etc. This collected data is stored in a centralized database.

[0685] Specific operation: The server calls the API to collect the latest user information, converts it into an appropriate format, and saves it in the database.

[0686] Step 2:

[0687] Training generative AI models and emotion engines

[0688] The server preprocesses the collected data and trains the generative AI model and emotion engine, and retrains the AI ​​model and emotion engine every time new data is added.

[0689] How it works: The server performs preprocessing such as data cleaning and tokenization, and uses a GPU cluster to train the AI ​​model and emotion engine.

[0690] Step 3:

[0691] Real-time data delivery

[0692] The server transmits the latest profile information and emotional state in real time in response to requests from the device.

[0693] Specific operation: The device sends a request to the server, and the server retrieves the user's information and emotional state from the database and returns it to the device in JSON format.

[0694] Step 4:

[0695] Face-to-face identification and emotion recognition

[0696] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, and then uses an emotion engine to recognize the other person's emotional state. It then queries the server, receives the other person's profile information and emotional state, and displays them.

[0697] How it works: The device captures the face of the person it's facing with a camera, runs a facial recognition algorithm and emotion engine, and sends the extracted facial features and emotion information to the server, which then returns the relevant information.

[0698] Step 5:

[0699] Voice recognition operation

[0700] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages. The voice data is sent to an emotion engine to analyze the user's emotions.

[0701] What it does: The device receives voice data, converts it to text using a speech recognition engine, executes an API request based on that information, and analyzes the voice emotion using an emotion engine.

[0702] Step 6:

[0703] Automating business operations with AI

[0704] The server uses AI generation after work hours or when the user is off-duty to automatically generate the optimal response based on the inquiry and the user's emotional state.

[0705] How it works: The server passes the query and sentiment data to the model, generates the best answer, and sends it back via email or messaging system.

[0706] Step 7:

[0707] Profile and sentiment review

[0708] Users can check the profile information and emotional state of the person they are meeting through the screen of their smart glasses or device, allowing for smooth communication.

[0709] Specific behavior: The user sees the profile and emotional state of the other person displayed on the smart glasses and responds or engages in conversation appropriately.

[0710] Step 8:

[0711] Control and confirm emotions with voice commands

[0712] Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and checking their own and the other person's emotional state.

[0713] Specific operation: The user issues a voice command, the device recognizes it, and performs the necessary operation, displaying emotional information analyzed by the emotion engine.

[0714] Step 9:

[0715] Reduced workload

[0716] The AI ​​can take over after work or while the user is off-duty, reducing the burden on the user. Feedback based on emotion recognition also reduces stress and enables appropriate work responses.

[0717] Specific operation: AI automatically generates and sends appropriate responses to inquiries. It analyzes emotional information and provides appropriate feedback to users.

[0718] Example 2

[0719] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0720] In recent years, changes in the work environment have made it increasingly difficult for employees to communicate smoothly. Furthermore, remote work and flexible office settings have reduced face-to-face communication, making it difficult to recognize emotions and streamline work. This can lead to delays in work progress and team cooperation. A system is needed to resolve these issues, promote smooth communication between employees, and improve work efficiency.

[0721] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0722] In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, an emotion engine means, a user profile collection means, a generative AI model training means, an emotion engine training means, a real-time data distribution means, a face-to-face person identification means, and an AI-based automation means for work responses. This enables smooth communication between employees and improves work efficiency.

[0723] "Facial recognition means" is a technology that uses cameras and sensors to identify a person's face.

[0724] "Data analysis methods using generative AI models" is a technology that utilizes AI models that use machine learning and deep learning to analyze collected data and generate useful information.

[0725] "Operation means using voice recognition" is a technology that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[0726] "Automatic response measures after work is completed" refers to technology that enables the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[0727] "Data transmission / reception means" refers to technology for transmitting and receiving data between different devices and servers inside and outside the system.

[0728] The "emotion engine means" is a technology for analyzing the user's voice and video data and recognizing the user's emotional state.

[0729] "Means for collecting user profiles" refers to technology that allows the server to collect data such as the user's name, department, work history, and email content from systems both inside and outside the company.

[0730] A "generative AI model training means" is a technique for training a generative AI model using collected data and for retraining the model whenever new data is added.

[0731] "Means for training the emotion engine" refers to a technology for training the emotion engine using stored audio and video data.

[0732] "Real-time data distribution means" refers to technology for transmitting the latest user profile information and emotional state in real time in response to a request from a terminal.

[0733] The "means for identifying the person being met" is a technology that recognizes the face of the person being met from camera images via smart glasses and queries the server.

[0734] "AI-based automation of business operations" is a technology that uses generative AI after work hours or during off-duty periods to automatically generate optimal responses based on inquiries and emotions.

[0735] This invention is an information analysis tool that recognizes user emotions by combining an emotion engine with "Smile Anywhere," thereby improving smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[0736] System Configuration

[0737] Facial Recognition Methods

[0738] The system uses cameras and sensors to capture the face of the person facing you and identify them. This method makes it possible to identify the other person even in situations where facial identification is difficult due to mask wearing or free address systems. Specifically, it utilizes face recognition technology (e.g., OpenCV) using camera images.

[0739] Data analysis methods using generative AI models

[0740] The system uses AI models (e.g., GPT-4) that use machine learning and deep learning to analyze the collected data and generate useful information, which can then present relevant information such as the other party's work history and email content.

[0741] Voice recognition operation

[0742] The system analyzes the user's voice, converts it into text, and operates the system according to the commands. This method allows hands-free employee search and message sending. Specifically, it uses a voice recognition service (e.g., Google Assistant).

[0743] Automated response measures after business hours

[0744] The system uses generative AI to automatically respond to inquiries after users have finished their work, giving users more freedom to use their time after work.

[0745] Data transmission and reception means

[0746] Send and receive data between different devices and servers inside and outside the system. This method makes it possible to obtain the necessary information in real time and display it on the appropriate device. Specifically, it uses RESTful APIs.

[0747] Emotion Engine

[0748] The system analyzes the user's voice and video data to recognize their emotional state, enabling it to respond and provide feedback according to their emotions.

[0749] How user profiles are collected

[0750] The server connects with internal and external systems to periodically collect data such as each user's name, department, work history, and email content, and centralizes it in a database. Specifically, it obtains data from LDAP directories and email servers via API.

[0751] A means of training generative AI models

[0752] The server uses the collected data to train a generative AI model, and continues to retrain the model as new data is added.

[0753] A means of training the emotion engine

[0754] The server trains the emotion engine based on the stored audio and video data. Specifically, it converts the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeds it into the emotion analysis model.

[0755] Real-time data delivery method

[0756] The server receives requests from the device and sends the latest user profile information and emotional state in real time. Specifically, it receives requests from the device using a RESTful API and returns them in JSON format.

[0757] Means of identifying the person you are facing

[0758] The device recognizes the face of the person in front of it from the camera image through the smart glasses, queries the server, and receives information from the server to display the other person's profile and emotional state. Specifically, it uses facial recognition software (e.g., OpenCV).

[0759] Automating business processes using AI

[0760] The device uses generative AI to automatically generate the optimal response based on the inquiry and the customer's emotion after work hours or when the device is in off mode. The AI ​​model is called via API, and the inquiry content and emotional data are input, and the generated response is automatically sent via email or chat.

[0761] Specific examples

[0762] For example, a user can recognize the face of the person they are meeting with during a meeting and display the person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, the AI ​​can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work. An example prompt might be, "Please describe a system that allows a user to identify the face of the person they are meeting with during a meeting and display the other person's emotional state and work history."

[0763] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0764] Step 1:

[0765] Collecting user profiles

[0766] The server connects with internal and external systems (for example, LDAP directories and mail servers) to collect data such as each user's name, department, work history, email content, etc. This data is obtained via API and centralized in the server's database.

[0767] Input: Data from an LDAP directory or mail server.

[0768] Output: Centralized database.

[0769] Step 2:

[0770] Training generative AI models

[0771] The server uses the collected data to train a generative AI model (e.g., GPT-4), and retrains this model to keep it up to date whenever new data is added.

[0772] Input: Information from a centralized database.

[0773] Output: A trained generative AI model.

[0774] Step 3:

[0775] Training the Emotional Engine

[0776] The server trains the emotion engine based on the stored audio and video data, for example by converting the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeding the text data into the emotion analysis model.

[0777] Input: Audio data, video data.

[0778] Output: A trained emotion engine.

[0779] Step 4:

[0780] Identifying the person you are facing

[0781] The device recognizes the face of the person it is meeting from the camera image through the smart glasses. Using facial recognition software (e.g., OpenCV), it sends the recognized facial data to the server, which then returns the other person's profile and emotional state.

[0782] Input: Camera footage.

[0783] Output: The other person's profile information and emotional state.

[0784] Step 5:

[0785] Viewing a User Profile

[0786] Users can check the profile information and emotional state of the person they are meeting through the smart glasses, which allows for smooth communication. The information is displayed on the smart glasses' HUD.

[0787] Input: Profile information and emotional state received from the server.

[0788] Output: Information displayed on the smart glasses HUD.

[0789] Step 6:

[0790] Voice recognition operation

[0791] The device uses a voice recognition function to analyze the user's voice commands. For example, it uses a voice recognition service (e.g., Google Assistant) to convert the voice commands into text and analyzes the text data. As a result, it becomes possible to perform operations such as searching for employees and sending messages.

[0792] Input: The user's voice command.

[0793] Output: Actions based on the parsed voice command.

[0794] Step 7:

[0795] Real-time data delivery

[0796] The server receives requests from the device and sends the latest user profile information and emotional state in real time. It uses a RESTful API to receive requests, retrieve the latest information from the database, and return it in JSON format.

[0797] Input: A request from the terminal.

[0798] Output: Latest user profile information and emotional state.

[0799] Step 8:

[0800] Automating business operations with AI

[0801] The device uses generative AI after work hours or when in off-duty mode to automatically generate the optimal response based on the inquiry and emotion. For example, the AI ​​model can be called via API, the inquiry content and emotion data are input, and the generated response is automatically sent via email or chat.

[0802] Input: Enquiry content and emotion data.

[0803] Output: The generated answer.

[0804] Step 9:

[0805] Reduced workload

[0806] The AI ​​will answer inquiries on behalf of the user after work or while the user is off-duty. It also reduces the burden of work by providing feedback based on emotion recognition. For example, the user can request, "Please send me an email with your thoughts on today's meeting," and the AI ​​will automatically respond.

[0807] Input: User request.

[0808] Output: Auto-generated emails and feedback.

[0809] (Application example 2)

[0810] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0811] Customer service in modern brick-and-mortar stores requires quick and effective responses to individual customers, but achieving this is not easy. For example, it is difficult for store employees to check inventory while serving a customer, or to appropriately understand and respond to a customer's emotional state. Furthermore, employees are required to respond quickly to customer inquiries even after their shifts have finished, placing a heavy burden on them. These challenges hinder efficient customer service and the improvement of business processes.

[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, a customer information display means using emotion recognition, and an inventory confirmation means using voice operation. This allows store employees to not only instantly check customer information and emotional state when meeting a customer, but also to quickly check inventory and suggest products using voice commands. Furthermore, realizing automatic response after work is completed using AI reduces the burden on employees and enables more efficient customer service.

[0813] "Facial recognition means" is a function that uses a camera or sensor to capture a person's face and identify a specific individual.

[0814] "Means for data analysis using generative AI models" refers to a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[0815] "Operation means using voice recognition" is a function that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[0816] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[0817] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[0818] The "means for displaying customer information based on emotion recognition" is a function that analyzes the customer's voice and video data, recognizes their emotional state, and displays that information.

[0819] The "voice-operated inventory check means" is a function that allows you to check the inventory status using voice commands.

[0820] System Overview

[0821] The present invention is a system for improving the efficiency of customer service in brick-and-mortar stores. This system includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after business hours are completed, a data transmission and reception unit, a customer information display unit using emotion recognition, and an inventory confirmation unit using voice operation.

[0822] Hardware and Software Use

[0823] The system uses the following hardware and software:

[0824] Hardware

[0825] Smart glasses (e.g., Google Glass, Vuzix Blade): Equipped with a camera and microphone to capture the face of the person you are facing.

[0826] High-performance camera: Captures video data for facial recognition.

[0827] Microphone: Captures voice commands and is used for voice recognition functions.

[0828] software

[0829] Operation OS: Android, iOS

[0830] Generative AI model: Data analysis using OpenAI, GPT-4

[0831] Emotion recognition engine: Affectiva

[0832] Speech Recognition API: Google Speech-to-Text

[0833] Cloud databases: Firebase, AWS RDS

[0834] Real-time data transfer: WebSocket, Firebase Realtime Database

[0835] Program processing

[0836] Acquiring customer information through facial recognition

[0837] The camera in the smart glasses captures the customer's face. The captured image data is analyzed by facial recognition and compared with a database to obtain the customer's profile information. This allows the necessary information to be displayed immediately when interacting with the customer.

[0838] Emotion-aware response

[0839] The captured video data is sent to an emotion recognition engine (Affectiva) to analyze the customer's emotional state. The resulting emotional state (e.g., satisfaction, confusion, interest, etc.) is displayed on the smart glasses' display, allowing store staff to quickly respond according to the customer's emotions.

[0840] Voice recognition operation

[0841] When an employee issues a voice command, the system uses the Google Speech-to-Text API to convert the voice into text, analyzes it, and performs the command accordingly. For example, if you issue the voice command "Check inventory, product number 12345," inventory information will be retrieved from a cloud database and displayed in real time.

[0842] Automatic response after business hours

[0843] Even after work hours, automated responses are provided using generative AI models (OpenAI, GPT-4). The AI ​​model can generate and send appropriate responses to customer inquiries, reducing the workload of employees.

[0844] Examples and prompts

[0845] Example 1: Checking inventory

[0846] When a store clerk issues a voice command through the smart glasses, such as "Check stock, product number 12345," the system retrieves the product's stock status in real time from the cloud database and displays it on the smart glasses' display.

[0847] Example 2: Customer service using sentiment analysis

[0848] An emotion recognition engine analyzes the video data captured by the smart glasses' camera, and if it indicates that the customer is confused, employees can immediately change their response.

[0849] ChatGPT prompt example

[0850] Prompt for generative AI models (OpenAI, GPT-4):

[0851] "The customer is confused. What should we do next?"

[0852] Customer profile information:

[0853] Name: Taro

[0854] Past purchase history: Many home appliances

[0855] Preferences: Products with the latest technology

[0856] Emotional state: Confused

[0857] Generation example:

[0858] "Taro, what kind of home appliances would you like to buy today? We also have the latest smart home devices."

[0859] In this way, the operation of the system can significantly improve the quality of customer service and business efficiency.

[0860] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0861] Step 1:

[0862] A user puts on the smart glasses and launches the application. The smart glasses' camera captures the customer's face, and the video data is input into the facial recognition means. The facial recognition model analyzes the video data and extracts facial feature points. The extracted feature points are sent to the server for matching with a database. The server searches the database, obtains the corresponding customer's profile information, and sends it to the terminal. The terminal displays this information on the smart glasses' display. This allows employees to check information such as the customer's name and purchase history in real time.

[0863] Step 2:

[0864] The customer's video data captured by the smart glasses is input into an emotion recognition engine (Affectiva). The emotion recognition engine analyzes the customer's facial expressions and identifies their emotional state. The identified emotional information (e.g., satisfaction, confusion, interest, etc.) is sent to the terminal and displayed on the smart glasses' display. This allows employees to quickly respond according to the customer's emotions.

[0865] Step 3:

[0866] When a user issues a voice command, the microphone in the smart glasses captures the voice. The captured voice data is input into a speech recognition API (Google Speech-to-Text) and converted into text. The converted text data is sent to the server, where it is analyzed by a generative AI model (OpenAI, GPT-4). For example, in response to the command "Check stock, product number 12345," the server searches a cloud database and retrieves stock information for product number 12345. This information is sent to the device and displayed on the smart glasses' display. This allows employees to respond quickly to customer questions without taking their eyes off the customer.

[0867] Step 4:

[0868] Even after the user has finished their work, the generative AI model (OpenAI, GPT-4) remains in standby mode on the server. When a customer makes an inquiry, the inquiry is entered into the server. The server uses the generative AI model to generate the optimal answer and sends it to the customer via an automated response system. This enables prompt customer responses even after hours, reducing the workload of employees.

[0869] Through these steps, the present invention is a system that can significantly improve the efficiency of customer service in physical stores and improve business processes.

[0870] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0871] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0872] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0873] [Third embodiment]

[0874] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0875] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0876] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0877] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0878] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0879] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0880] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0881] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0882] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0883] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0884] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0885] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0886] The present invention, "Dokodemo Smile," is an information analysis tool that solves communication issues that arise due to changes in the working environment, and is composed of the following elements.

[0887] System Configuration

[0888] 1. Facial Recognition Methods

[0889] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[0890] 2. Data Analysis Methods Using Generative AI Models

[0891] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[0892] 3. Voice recognition operation

[0893] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[0894] 4. Automated response measures after business hours

[0895] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[0896] 5. Data transmission and reception means

[0897] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[0898] Program processing

[0899] The program processing of this system will be specifically explained below.

[0900] 1. Server Processing

[0901] Collection of user profiles: The server periodically collects data such as each user's name, department, work history, email content, etc. in cooperation with internal and external systems. This collected data is stored in a database.

[0902] Training the generative AI model: The server trains the generative AI model using the collected data and continues to retrain the model as new data is added.

[0903] Real-time data delivery: The server sends the latest profile information to the device in real time whenever a request is made.

[0904] 2. Terminal processing

[0905] Identifying the person in front of you: The device recognizes the face of the person in front of you from the camera image through the smart glasses and queries the server. It receives information from the server and displays the person's profile.

[0906] Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[0907] Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to generate optimal responses to inquiries and respond automatically.

[0908] 3. User Processing

[0909] Profile confirmation: Users can check the profile information of the person they are meeting through the screen of their smart glasses or device, enabling smooth communication.

[0910] Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, and sending messages.

[0911] Reduced workload: Users can reduce their workload by having AI respond to inquiries on their behalf after work or when they are off-duty.

[0912] Specific examples

[0913] For example, during a meeting, the smart glasses can recognize the face of the person they are meeting and display the person's name, department, and past email content. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[0914] summary

[0915] This invention, "Dokodemo Smile," is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automatic response after work, and data transmission and reception, in order to solve communication issues caused by changes in the work environment. This will realize smooth communication and improved work efficiency.

[0916] The processing flow will be explained below.

[0917] Step 1:

[0918] Collecting user profiles

[0919] The server connects with internal and external systems and periodically collects data such as each user's name, department, work history, email content, etc. This collected data is stored centrally in the server's database.

[0920] What happens: The server calls the API to collect the latest user information, converts it into an appropriate format, and stores it in the database.

[0921] Step 2:

[0922] Training generative AI models

[0923] The server preprocesses the collected data and trains generative AI models, and also retrains existing models as new data is added.

[0924] Specific operation: The server performs preprocessing such as data cleaning and tokenization, and then trains the AI ​​model using a GPU cluster.

[0925] Step 3:

[0926] Real-time data delivery

[0927] The server sends the latest profile information in real time in response to requests from the device.

[0928] Specific operation: The device sends a request to the server, and the server retrieves the user's information from the database and returns it to the device in JSON format.

[0929] Step 4:

[0930] Identifying the person you are facing

[0931] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, queries the server, receives information from the server, and displays the other person's profile.

[0932] What it does: The device runs a facial recognition algorithm, extracts facial features, and sends them to the server, which then returns the information.

[0933] Step 5:

[0934] Voice recognition operation

[0935] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages.

[0936] What it does: The device receives voice data, converts it to text using a speech recognition engine, and executes API requests based on that information.

[0937] Step 6:

[0938] Automating business operations with AI

[0939] The server uses the generation AI even after business hours or when the server is off-line to automatically generate optimal responses to inquiries.

[0940] How it works: The server passes the query to the AI ​​model, which generates an appropriate answer and replies via email or messaging system.

[0941] Step 7:

[0942] Profile confirmation

[0943] Users can check the profile information of the person they are meeting on the smart glasses display, enabling smooth communication.

[0944] Specific operation: The user decides what to talk about and what questions to ask based on the information displayed on the smart glasses.

[0945] Step 8:

[0946] Control with voice commands

[0947] Users can issue voice commands to the device to search for employees, make calls, send messages, and more.

[0948] Specific operation: The device recognizes the voice command given by the user and performs the appropriate operation.

[0949] Step 9:

[0950] Reduced workload

[0951] Users can reduce their workload by having AI take over after work or when they are off-duty.

[0952] Specific behavior: AI automatically responds to inquiries, reducing the tasks that users have to perform.

[0953] Example 1

[0954] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0955] As the work environment changes, communication between employees is becoming less smooth. The constant wearing of masks and the introduction of a free address system can make facial recognition difficult, making face-to-face information sharing inconvenient. Furthermore, if work continues uninterrupted after the end of the working day, there is also the problem of infringing on employees' private time.

[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0957] In this invention, the server includes a face recognition means, a data analysis means using a machine learning model, an operation means for analyzing voice input, an automatic response means after work hours, and an information transmission and reception means. This allows employees to easily identify the person they are meeting and efficiently operate using voice commands. In addition, since the AI ​​automatically responds after work hours, employees can protect their private time.

[0958] "Facial identification means" refers to a technical means for identifying the face of the target person using video equipment.

[0959] "Means for data analysis using machine learning models" refers to technical means for performing data analysis using machine learning algorithms based on collected data.

[0960] "Operation means for analyzing voice input" refers to a technical means for analyzing a user's voice, converting it into text, and operating the system according to the commands.

[0961] "Automatic response measures after work is completed" refers to a technical measure that allows the generation AI to automatically respond appropriately to inquiries even after the user's work is completed.

[0962] "Means for transmitting and receiving information" refers to the technical means for transmitting and receiving information between different devices and servers inside and outside the system.

[0963] "Video equipment" is a general term for hardware devices used to capture images, such as cameras and sensors.

[0964] A "machine learning algorithm" is a mathematical model that learns patterns from large amounts of data and automates tasks such as prediction and classification.

[0965] A "voice command" is an instruction to operate the system using voice.

[0966] The present invention, "Dokodemo Smile," is an information analysis system that combines multiple technical means to solve communication issues that arise with changes in the work environment. The present invention is composed of the following main components:

[0967] Facial Identification Method

[0968] The device has the ability to identify the face of the person in front of it using video equipment such as cameras and sensors. Facial identification uses OpenCV and Microsoft Azure's Face API. This method can eliminate the difficulty of face identification due to mask wearing and free address systems. For example, during a meeting, the camera in the smart glasses can detect the face of the person in front of it, convert the facial features into vectors, and send them to a server.

[0969] Data analysis methods using machine learning models

[0970] The server analyzes the collected data using machine learning algorithms (e.g., GPT-3 or BERT). This data includes the user's name, department, work history, email content, etc., and is updated regularly. The analyzed information is sent to the device in real time whenever a request is made. For example, if a user wants to check the profile of a person they are meeting with, the server will provide the latest profile information.

[0971] A means of operation that analyzes voice input

[0972] The device analyzes the user's voice, converts it into text, and operates the system according to the commands. This voice recognition function uses Google Cloud Speech-to-Text and Amazon Lex. Users can use voice commands to perform operations such as searching for employees, making calls, and sending messages. For example, if you issue the voice command "View Tanaka's profile," the voice will be converted into text and the device will perform the corresponding operation.

[0973] Automated response measures after business hours

[0974] The device uses generative AI (e.g., GPT-3 or BERT) to generate optimal responses to inquiries after work hours or while in off-mode, and automatically responds. This allows users to spend their time with peace of mind after work. For example, if a user receives an inquiry email asking, "Please tell me more about tomorrow's meeting," while in off-mode, the AI ​​will automatically generate and send a reply.

[0975] Means of sending and receiving information

[0976] The server and terminals send and receive information between different devices and servers inside and outside the system. For this purpose, a real-time data streaming platform such as Apache Kafka is used. This makes it possible to obtain the necessary information in real time and display it on the terminal. For example, if new profile data is sent from the server during a meeting, it will be displayed on the terminal immediately.

[0977] Specific examples

[0978] For example, during a meeting, a user can recognize the face of the person they are meeting and display the person's name, department, and past email content on the smart glasses. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[0979] Examples of prompt statements

[0980] "Facial recognition is performed during meetings, and the other person's name, department, and past email content are displayed. The conversation proceeds smoothly, and documents can be emailed with subsequent voice commands. AI can automatically respond to inquiries even after work has finished."

[0981] In this way, the present invention, "Smile Anywhere," integrates a variety of technological means and responds to changes in the work environment, thereby realizing smooth communication and improved work efficiency.

[0982] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0983] Program processing flow and specific explanation

[0984] Step 1: Collecting user profiles and storing them in a database

[0985] The server periodically collects data such as user name, department, work history, email content, etc. by linking with internal and external systems. This linking is done using REST API and SOAP.

[0986] Input: User name, department, work history, email content, etc. obtained through API calls

[0987] Data processing: Analyzes acquired data in JSON or XML format and converts it into a format that can be stored in a database.

[0988] Output: Correctly formatted user data stored in the database

[0989] What it does: Every night, the server uses a job scheduler (e.g., cron) to send a request to a specified API endpoint to retrieve the latest user data, which is then parsed and stored in a database.

[0990] Step 2: Train and update the generative AI model

[0991] The server uses the collected data to train a generative AI (e.g., GPT-3, BERT) and updates the model whenever new data is added.

[0992] Input: Latest user data stored in the database

[0993] Data computation: Using machine learning algorithms to train models and optimize parameters

[0994] Output: An updated generative AI model

[0995] How it works: Every weekend, the server begins retraining the model using newly collected data. It uses a machine learning framework (e.g., TensorFlow or PyTorch) to learn patterns from large amounts of data and improve the model's accuracy.

[0996] Step 3: Identify your contact and view their profile

[0997] The device recognizes the face of the person it is meeting from the camera image via the smart glasses, queries the server and displays their profile.

[0998] Input: Smartglasses camera image

[0999] Data processing: Detecting face regions from video and generating face feature vectors

[1000] Output: Send the facial feature vector to the server, retrieve the corresponding person's profile information, and display it on the smart glasses.

[1001] How it works: During a meeting, the smart glasses analyze the camera footage in real time, use a facial recognition API to identify the face of the person they are meeting with, query the server to obtain the person's profile information, and display it on the smart glasses' display.

[1002] Step 4: Control with voice input

[1003] The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[1004] Input: User's voice

[1005] Data processing: Converting voice to text and parsing it as commands

[1006] Output: Based on the interpreted command, perform the corresponding operation on the terminal.

[1007] What happens: When a user says "View Tanaka's profile," the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the speech to text and take the appropriate action.

[1008] Step 5: Automated responses after the business day ends

[1009] The device responds to inquiries after work hours or in off-mode using the generative AI model, generating optimal answers and responding automatically.

[1010] Input: User inquiry (e.g., email)

[1011] Data Computation: Using generative AI models to generate relevant answers

[1012] Output: Send the generated answer to the user

[1013] Specific operation: When a user receives an inquiry email after work asking, "Please tell me the details about tomorrow's meeting," the device uses the generative AI model to automatically generate and send a response saying, "Tomorrow's meeting will be held in conference room A from 10:00."

[1014] Step 6: Send and receive information in real time

[1015] The server and terminal use a real-time data streaming platform such as Apache Kafka to send and receive information in real time.

[1016] Input: Newly acquired data and updated information

[1017] Data processing: process data in real time and send it to the device in the appropriate format

[1018] Output: Real-time updated information is displayed on the terminal.

[1019] Specific operation: When new profile data is sent from the server during a meeting, it is immediately displayed on the device, allowing the user to view the latest information.

[1020] In this way, the present invention "Smile Anywhere" smoothly solves communication issues that arise with changes in the work environment through multiple processing steps.

[1021] (Application example 1)

[1022] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1023] In recent years, there has been a demand for personalized service for each customer in brick-and-mortar stores. However, it is difficult to instantly grasp a customer's name, face, and purchase history, and there is a lack of systems to ensure smooth customer service. In addition, staff are required to respond to customer inquiries even after work hours, which increases the burden on staff. There is a need for an efficient customer service support system to solve these issues.

[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1025] In this invention, the server includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after work is completed, a data transmission and reception unit, a unit for acquiring customer profile information, and a unit for analyzing purchase history. This enables personalized responses to each customer in a physical store, reducing the workload of staff and improving customer satisfaction.

[1026] "Facial recognition means" is a function that uses a camera or sensor to capture and identify a person's face.

[1027] "Data analysis means using generative AI models" is a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[1028] "Operation means using voice recognition" is a function that analyzes voice commands, converts them into text, and operates the system according to those commands.

[1029] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[1030] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[1031] "Means for obtaining customer profile information" refers to a function that collects data such as customer name, affiliation, purchase history, and survey responses, and stores it in a database.

[1032] "Means for analyzing purchasing history" refers to a function that analyzes a customer's past purchasing history and understands their preferences and trends.

[1033] The following detailed description of the embodiments of the present invention will be given. Note that the embodiments described herein embody the technical features included in the claims, but do not limit the technical scope of the invention.

[1034] A customer support system in a physical store mainly consists of three components: a server, a terminal, and a user. The roles and operations of each component are as follows:

[1035] 1. Server Roles and Operations

[1036] The server plays a central role in acquiring customer profile information and analyzing the data using generative AI models. The server operates using the following hardware and software:

[1037] Hardware: Server equipment with high-performance processors, memory, and large storage capacity.

[1038] Software: Database management systems (e.g., MySQL), deep learning frameworks (e.g., TensorFlow), speech recognition APIs (e.g., Google Speech-to-Text).

[1039] The server performs the following process:

[1040] Collecting customer profiles: The server periodically collects data such as customer names, purchase history, and survey responses and stores it in a database.

[1041] Training the generative AI model: The server trains the generative AI model using the collected data and retrains the model whenever new data is added.

[1042] Real-time data delivery: The server sends the latest customer profile information to the device in real time upon request.

[1043] 2. Roles and Functions of the Device

[1044] Terminals are devices used by customer-facing staff, such as smart glasses and tablets, that operate using the following hardware and software:

[1045] Hardware: Smart glasses, camera, microphone.

[1046] Software: Facial recognition libraries (e.g., OpenCV), speech recognition software (e.g., Google Speech-to-Text).

[1047] The terminal performs the following process:

[1048] Customer face identification: The device recognizes the customer's face from the camera image of the smart glasses and queries the server for that information. The customer profile information sent from the server is displayed on the smart glasses display.

[1049] Voice recognition operation: The terminal recognizes voice commands and performs operations such as displaying customer information, checking inventory, and operating the cash register.

[1050] Automatic response after business hours: The device will use generative AI to automatically respond to inquiries after business hours or while in off mode.

[1051] 3. User Roles and Actions

[1052] The users of this system are the service staff who check customer information through the smart glasses or the terminal screen and perform operations using voice commands.

[1053] Check customer information: Staff can view information such as customer profiles and purchase history through the smart glasses screen.

[1054] Voice command operation: Staff use voice commands to search for information, check inventory, operate the cash register, and more.

[1055] Reduced workload: After the work is completed, AI automatically handles customer support, reducing the workload of staff.

[1056] Specific examples

[1057] For example, a sales staff member at a physical store can wear smart glasses and recognize the face of a customer who visits the store. The system displays the customer's profile and purchase history on the smart glasses' display, allowing the staff member to recommend products based on the customer's preferences and past purchase history. Furthermore, when the staff member issues a voice command such as "Show me recommended products for this customer," the system uses AI to select and display the appropriate products.

[1058] This will enable personalized service in physical stores, improving customer satisfaction. In addition, the AI ​​will automatically respond to customers after the store has finished its work, reducing the workload of staff.

[1059] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1060] Step 1:

[1061] The server collects customer profile information and stores it in a database. Specifically, it periodically obtains customer names, purchase histories, survey responses, etc. from the POS system in the physical store and the online survey system. The input to this process is data from the POS system and the survey system, and the output is customer information stored in the database.

[1062] Step 2:

[1063] The server uses the stored customer information to train a generative AI model. Specifically, it uses a deep learning framework (e.g., TensorFlow) to analyze customer purchasing trends and preferences. The input to this process is the customer information in the database, and the output is the trained AI model.

[1064] Step 3:

[1065] The server receives requests from the smart glasses and delivers real-time customer profile information. Specifically, it receives facial recognition data sent from the smart glasses' camera, retrieves the corresponding customer information from the database, and sends it to the smart glasses. The input of this process is facial recognition data, and the output is customer profile information.

[1066] Step 4:

[1067] The device (smart glasses) recognizes the customer's face using camera images. Specifically, it uses a facial recognition library (e.g., OpenCV) to extract facial features from the image and sends that information to a server. The input to this process is the camera image, and the output is facial recognition data.

[1068] Step 5:

[1069] The terminal recognizes the voice command and performs the required action. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text) to convert the voice command into text and follows the instructions to display customer information or check inventory. The input to this process is the voice command, and the output is the text command and the result of the action taken.

[1070] Step 6:

[1071] The terminal automatically responds to inquiries even after work hours. Specifically, it uses generative AI to generate appropriate responses to customer inquiries and automatically replies. The input to this process is the customer inquiry, and the output is the generated response.

[1072] Step 7:

[1073] The user checks customer information through the smart glasses screen and uses voice commands. Specifically, the user reads the information displayed on the smart glasses display and takes the necessary action to respond to the customer. The input of this process is the customer information displayed on the display, and the output is the user's specific action.

[1074] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1075] This invention is an information analysis tool that combines "Smile Anywhere" with an emotion engine to recognize user emotions and further improve smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[1076] System Configuration

[1077] 1. Facial Recognition Methods

[1078] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[1079] 2. Data Analysis Methods Using Generative AI Models

[1080] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[1081] 3. Voice recognition operation

[1082] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[1083] 4. Automated response measures after business hours

[1084] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[1085] 5. Data transmission and reception means

[1086] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[1087] 6. Emotion Engine

[1088] This function analyzes the user's voice and video data to recognize their emotional state, enabling them to respond and provide feedback according to their emotions.

[1089] Program processing

[1090] The system program of the present invention performs the following processes.

[1091] 1. Server Processing

[1092] - User profile collection: The server periodically collects data such as each user's name, department, work history, and email content in cooperation with internal and external systems. This data is then centralized in the server's database.

[1093] - Training generative AI models: The server trains generative AI models using collected data and continues to retrain existing models as new data is added.

[1094] - Training the emotion engine: The server trains the emotion engine based on the accumulated audio and video data.

[1095] - Real-time data delivery: The server sends the latest profile information and emotional state in real time in response to requests from the device.

[1096] 2. Terminal processing

[1097] - Face-to-face identification: The device recognizes the face of the person it is facing from the camera image through the smart glasses and queries the server. It receives information from the server and displays the face's profile and emotional state.

[1098] - Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees and sending messages. It also sends the voice data to an emotion engine for emotion recognition.

[1099] - Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to automatically generate the optimal response based on the inquiry and emotion.

[1100] 3. User Processing

[1101] - Profile Check: Users can check the profile information and emotional state of the person they are meeting through the smart glasses or device screen, enabling smooth communication.

[1102] - Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and also check the results of voice emotion analysis.

[1103] - Reduced workload: After work or while users are off-duty, AI can answer inquiries on their behalf and provide feedback based on emotion recognition, reducing their workload.

[1104] Specific examples

[1105] For example, during a meeting, a user can recognize the face of the person they are meeting and display the other person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, AI can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work.

[1106] summary

[1107] This invention is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automated responses after work is completed, data transmission and reception, and an emotion engine, to resolve communication barriers caused by changes in the work environment, thereby achieving smoother and more efficient communication and improved work efficiency.

[1108] The processing flow will be explained below.

[1109] Step 1:

[1110] Collecting user profiles

[1111] The server connects with internal and external systems and periodically collects each user's name, department, work history, email content, voice data, etc. This collected data is stored in a centralized database.

[1112] Specific operation: The server calls the API to collect the latest user information, converts it into an appropriate format, and saves it in the database.

[1113] Step 2:

[1114] Training generative AI models and emotion engines

[1115] The server preprocesses the collected data and trains the generative AI model and emotion engine, and retrains the AI ​​model and emotion engine every time new data is added.

[1116] How it works: The server performs preprocessing such as data cleaning and tokenization, and uses a GPU cluster to train the AI ​​model and emotion engine.

[1117] Step 3:

[1118] Real-time data delivery

[1119] The server transmits the latest profile information and emotional state in real time in response to requests from the device.

[1120] Specific operation: The device sends a request to the server, and the server retrieves the user's information and emotional state from the database and returns it to the device in JSON format.

[1121] Step 4:

[1122] Face-to-face identification and emotion recognition

[1123] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, and then uses an emotion engine to recognize the other person's emotional state. It then queries the server, receives the other person's profile information and emotional state, and displays them.

[1124] How it works: The device captures the face of the person it's facing with a camera, runs a facial recognition algorithm and emotion engine, and sends the extracted facial features and emotion information to the server, which then returns the relevant information.

[1125] Step 5:

[1126] Voice recognition operation

[1127] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages. The voice data is sent to an emotion engine to analyze the user's emotions.

[1128] What it does: The device receives voice data, converts it to text using a speech recognition engine, executes an API request based on that information, and analyzes the voice emotion using an emotion engine.

[1129] Step 6:

[1130] Automating business operations with AI

[1131] The server uses AI generation after work hours or when the user is off-duty to automatically generate the optimal response based on the inquiry and the user's emotional state.

[1132] How it works: The server passes the query and sentiment data to the model, generates the best answer, and sends it back via email or messaging system.

[1133] Step 7:

[1134] Profile and sentiment review

[1135] Users can check the profile information and emotional state of the person they are meeting through the screen of their smart glasses or device, allowing for smooth communication.

[1136] Specific behavior: The user sees the profile and emotional state of the other person displayed on the smart glasses and responds or engages in conversation appropriately.

[1137] Step 8:

[1138] Control and confirm emotions with voice commands

[1139] Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and checking their own and the other person's emotional state.

[1140] Specific operation: The user issues a voice command, the device recognizes it, and performs the necessary operation, displaying emotional information analyzed by the emotion engine.

[1141] Step 9:

[1142] Reduced workload

[1143] The AI ​​can take over after work or while the user is off-duty, reducing the burden on the user. Feedback based on emotion recognition also reduces stress and enables appropriate work responses.

[1144] Specific operation: AI automatically generates and sends appropriate responses to inquiries. It analyzes emotional information and provides appropriate feedback to users.

[1145] Example 2

[1146] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1147] In recent years, changes in the work environment have made it increasingly difficult for employees to communicate smoothly. Furthermore, remote work and flexible office settings have reduced face-to-face communication, making it difficult to recognize emotions and streamline work. This can lead to delays in work progress and team cooperation. A system is needed to resolve these issues, promote smooth communication between employees, and improve work efficiency.

[1148] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1149] In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, an emotion engine means, a user profile collection means, a generative AI model training means, an emotion engine training means, a real-time data distribution means, a face-to-face person identification means, and an AI-based automation means for work responses. This enables smooth communication between employees and improves work efficiency.

[1150] "Facial recognition means" is a technology that uses cameras and sensors to identify a person's face.

[1151] "Data analysis methods using generative AI models" is a technology that utilizes AI models that use machine learning and deep learning to analyze collected data and generate useful information.

[1152] "Operation means using voice recognition" is a technology that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[1153] "Automatic response measures after work is completed" refers to technology that enables the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[1154] "Data transmission / reception means" refers to technology for transmitting and receiving data between different devices and servers inside and outside the system.

[1155] The "emotion engine means" is a technology for analyzing the user's voice and video data and recognizing the user's emotional state.

[1156] "Means for collecting user profiles" refers to technology that allows the server to collect data such as the user's name, department, work history, and email content from systems both inside and outside the company.

[1157] A "generative AI model training means" is a technique for training a generative AI model using collected data and for retraining the model whenever new data is added.

[1158] "Means for training the emotion engine" refers to a technology for training the emotion engine using stored audio and video data.

[1159] "Real-time data distribution means" refers to technology for transmitting the latest user profile information and emotional state in real time in response to a request from a terminal.

[1160] The "means for identifying the person being met" is a technology that recognizes the face of the person being met from camera images via smart glasses and queries the server.

[1161] "AI-based automation of business operations" is a technology that uses generative AI after work hours or during off-duty periods to automatically generate optimal responses based on inquiries and emotions.

[1162] This invention is an information analysis tool that recognizes user emotions by combining an emotion engine with "Smile Anywhere," thereby improving smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[1163] System Configuration

[1164] Facial Recognition Methods

[1165] The system uses cameras and sensors to capture the face of the person facing you and identify them. This method makes it possible to identify the other person even in situations where facial identification is difficult due to mask wearing or free address systems. Specifically, it utilizes face recognition technology (e.g., OpenCV) using camera images.

[1166] Data analysis methods using generative AI models

[1167] The system uses AI models (e.g., GPT-4) that use machine learning and deep learning to analyze the collected data and generate useful information, which can then present relevant information such as the other party's work history and email content.

[1168] Voice recognition operation

[1169] The system analyzes the user's voice, converts it into text, and operates the system according to the commands. This method allows hands-free employee search and message sending. Specifically, it uses a voice recognition service (e.g., Google Assistant).

[1170] Automated response measures after business hours

[1171] The system uses generative AI to automatically respond to inquiries after users have finished their work, giving users more freedom to use their time after work.

[1172] Data transmission and reception means

[1173] Send and receive data between different devices and servers inside and outside the system. This method makes it possible to obtain the necessary information in real time and display it on the appropriate device. Specifically, it uses RESTful APIs.

[1174] Emotion Engine

[1175] The system analyzes the user's voice and video data to recognize their emotional state, enabling it to respond and provide feedback according to their emotions.

[1176] How user profiles are collected

[1177] The server connects with internal and external systems to periodically collect data such as each user's name, department, work history, and email content, and centralizes it in a database. Specifically, it obtains data from LDAP directories and email servers via API.

[1178] A means of training generative AI models

[1179] The server uses the collected data to train a generative AI model, and continues to retrain the model as new data is added.

[1180] A means of training the emotion engine

[1181] The server trains the emotion engine based on the stored audio and video data. Specifically, it converts the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeds it into the emotion analysis model.

[1182] Real-time data delivery method

[1183] The server receives requests from the device and sends the latest user profile information and emotional state in real time. Specifically, it receives requests from the device using a RESTful API and returns them in JSON format.

[1184] Means of identifying the person you are facing

[1185] The device recognizes the face of the person in front of it from the camera image through the smart glasses, queries the server, and receives information from the server to display the other person's profile and emotional state. Specifically, it uses facial recognition software (e.g., OpenCV).

[1186] Automating business processes using AI

[1187] The device uses generative AI to automatically generate the optimal response based on the inquiry and the customer's emotion after work hours or when the device is in off mode. The AI ​​model is called via API, and the inquiry content and emotional data are input, and the generated response is automatically sent via email or chat.

[1188] Specific examples

[1189] For example, a user can recognize the face of the person they are meeting with during a meeting and display the person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, the AI ​​can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work. An example prompt might be, "Please describe a system that allows a user to identify the face of the person they are meeting with during a meeting and display the other person's emotional state and work history."

[1190] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1191] Step 1:

[1192] Collecting user profiles

[1193] The server connects with internal and external systems (for example, LDAP directories and mail servers) to collect data such as each user's name, department, work history, email content, etc. This data is obtained via API and centralized in the server's database.

[1194] Input: Data from an LDAP directory or mail server.

[1195] Output: Centralized database.

[1196] Step 2:

[1197] Training generative AI models

[1198] The server uses the collected data to train a generative AI model (e.g., GPT-4), and retrains this model to keep it up to date whenever new data is added.

[1199] Input: Information from a centralized database.

[1200] Output: A trained generative AI model.

[1201] Step 3:

[1202] Training the Emotional Engine

[1203] The server trains the emotion engine based on the stored audio and video data, for example by converting the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeding the text data into the emotion analysis model.

[1204] Input: Audio data, video data.

[1205] Output: A trained emotion engine.

[1206] Step 4:

[1207] Identifying the person you are facing

[1208] The device recognizes the face of the person it is meeting from the camera image through the smart glasses. Using facial recognition software (e.g., OpenCV), it sends the recognized facial data to the server, which then returns the other person's profile and emotional state.

[1209] Input: Camera footage.

[1210] Output: The other person's profile information and emotional state.

[1211] Step 5:

[1212] Viewing a User Profile

[1213] Users can check the profile information and emotional state of the person they are meeting through the smart glasses, which allows for smooth communication. The information is displayed on the smart glasses' HUD.

[1214] Input: Profile information and emotional state received from the server.

[1215] Output: Information displayed on the smart glasses HUD.

[1216] Step 6:

[1217] Voice recognition operation

[1218] The device uses a voice recognition function to analyze the user's voice commands. For example, it uses a voice recognition service (e.g., Google Assistant) to convert the voice commands into text and analyzes the text data. As a result, it becomes possible to perform operations such as searching for employees and sending messages.

[1219] Input: The user's voice command.

[1220] Output: Actions based on the parsed voice command.

[1221] Step 7:

[1222] Real-time data delivery

[1223] The server receives requests from the device and sends the latest user profile information and emotional state in real time. It uses a RESTful API to receive requests, retrieve the latest information from the database, and return it in JSON format.

[1224] Input: A request from the terminal.

[1225] Output: Latest user profile information and emotional state.

[1226] Step 8:

[1227] Automating business operations with AI

[1228] The device uses generative AI after work hours or when in off-duty mode to automatically generate the optimal response based on the inquiry and emotion. For example, the AI ​​model can be called via API, the inquiry content and emotion data are input, and the generated response is automatically sent via email or chat.

[1229] Input: Enquiry content and emotion data.

[1230] Output: The generated answer.

[1231] Step 9:

[1232] Reduced workload

[1233] The AI ​​will answer inquiries on behalf of the user after work or while the user is off-duty. It also reduces the burden of work by providing feedback based on emotion recognition. For example, the user can request, "Please send me an email with your thoughts on today's meeting," and the AI ​​will automatically respond.

[1234] Input: User request.

[1235] Output: Auto-generated emails and feedback.

[1236] (Application example 2)

[1237] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1238] Customer service in modern brick-and-mortar stores requires quick and effective responses to individual customers, but achieving this is not easy. For example, it is difficult for store employees to check inventory while serving a customer, or to appropriately understand and respond to a customer's emotional state. Furthermore, employees are required to respond quickly to customer inquiries even after their shifts have finished, placing a heavy burden on them. These challenges hinder efficient customer service and the improvement of business processes.

[1239] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, a customer information display means using emotion recognition, and an inventory confirmation means using voice operation. This allows store employees to not only instantly check customer information and emotional state when meeting a customer, but also to quickly check inventory and suggest products using voice commands. Furthermore, realizing automatic response after work is completed using AI reduces the burden on employees and enables more efficient customer service.

[1240] "Facial recognition means" is a function that uses a camera or sensor to capture a person's face and identify a specific individual.

[1241] "Means for data analysis using generative AI models" refers to a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[1242] "Operation means using voice recognition" is a function that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[1243] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[1244] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[1245] The "means for displaying customer information based on emotion recognition" is a function that analyzes the customer's voice and video data, recognizes their emotional state, and displays that information.

[1246] The "voice-operated inventory check means" is a function that allows you to check the inventory status using voice commands.

[1247] System Overview

[1248] The present invention is a system for improving the efficiency of customer service in brick-and-mortar stores. This system includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after business hours are completed, a data transmission and reception unit, a customer information display unit using emotion recognition, and an inventory confirmation unit using voice operation.

[1249] Hardware and Software Use

[1250] The system uses the following hardware and software:

[1251] Hardware

[1252] Smart glasses (e.g., Google Glass, Vuzix Blade): Equipped with a camera and microphone to capture the face of the person you are facing.

[1253] High-performance camera: Captures video data for facial recognition.

[1254] Microphone: Captures voice commands and is used for voice recognition functions.

[1255] software

[1256] Operation OS: Android, iOS

[1257] Generative AI model: Data analysis using OpenAI, GPT-4

[1258] Emotion recognition engine: Affectiva

[1259] Speech Recognition API: Google Speech-to-Text

[1260] Cloud databases: Firebase, AWS RDS

[1261] Real-time data transfer: WebSocket, Firebase Realtime Database

[1262] Program processing

[1263] Acquiring customer information through facial recognition

[1264] The camera in the smart glasses captures the customer's face. The captured image data is analyzed by facial recognition and compared with a database to obtain the customer's profile information. This allows the necessary information to be displayed immediately when interacting with the customer.

[1265] Emotion-aware response

[1266] The captured video data is sent to an emotion recognition engine (Affectiva) to analyze the customer's emotional state. The resulting emotional state (e.g., satisfaction, confusion, interest, etc.) is displayed on the smart glasses' display, allowing store staff to quickly respond according to the customer's emotions.

[1267] Voice recognition operation

[1268] When an employee issues a voice command, the system uses the Google Speech-to-Text API to convert the voice into text, analyzes it, and performs the command accordingly. For example, if you issue the voice command "Check inventory, product number 12345," inventory information will be retrieved from a cloud database and displayed in real time.

[1269] Automatic response after business hours

[1270] Even after work hours, automated responses are provided using generative AI models (OpenAI, GPT-4). The AI ​​model can generate and send appropriate responses to customer inquiries, reducing the workload of employees.

[1271] Examples and prompts

[1272] Example 1: Checking inventory

[1273] When a store clerk issues a voice command through the smart glasses, such as "Check stock, product number 12345," the system retrieves the product's stock status in real time from the cloud database and displays it on the smart glasses' display.

[1274] Example 2: Customer service using sentiment analysis

[1275] An emotion recognition engine analyzes the video data captured by the smart glasses' camera, and if it indicates that the customer is confused, employees can immediately change their response.

[1276] ChatGPT prompt example

[1277] Prompt for generative AI models (OpenAI, GPT-4):

[1278] "The customer is confused. What should we do next?"

[1279] Customer profile information:

[1280] Name: Taro

[1281] Past purchase history: Many home appliances

[1282] Preferences: Products with the latest technology

[1283] Emotional state: Confused

[1284] Generation example:

[1285] "Taro, what kind of home appliances would you like to buy today? We also have the latest smart home devices."

[1286] In this way, the operation of the system can significantly improve the quality of customer service and business efficiency.

[1287] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1288] Step 1:

[1289] A user puts on the smart glasses and launches the application. The smart glasses' camera captures the customer's face, and the video data is input into the facial recognition means. The facial recognition model analyzes the video data and extracts facial feature points. The extracted feature points are sent to the server for matching with a database. The server searches the database, obtains the corresponding customer's profile information, and sends it to the terminal. The terminal displays this information on the smart glasses' display. This allows employees to check information such as the customer's name and purchase history in real time.

[1290] Step 2:

[1291] The customer's video data captured by the smart glasses is input into an emotion recognition engine (Affectiva). The emotion recognition engine analyzes the customer's facial expressions and identifies their emotional state. The identified emotional information (e.g., satisfaction, confusion, interest, etc.) is sent to the terminal and displayed on the smart glasses' display. This allows employees to quickly respond according to the customer's emotions.

[1292] Step 3:

[1293] When a user issues a voice command, the microphone in the smart glasses captures the voice. The captured voice data is input into a speech recognition API (Google Speech-to-Text) and converted into text. The converted text data is sent to the server, where it is analyzed by a generative AI model (OpenAI, GPT-4). For example, in response to the command "Check stock, product number 12345," the server searches a cloud database and retrieves stock information for product number 12345. This information is sent to the device and displayed on the smart glasses' display. This allows employees to respond quickly to customer questions without taking their eyes off the customer.

[1294] Step 4:

[1295] Even after the user has finished their work, the generative AI model (OpenAI, GPT-4) remains in standby mode on the server. When a customer makes an inquiry, the inquiry is entered into the server. The server uses the generative AI model to generate the optimal answer and sends it to the customer via an automated response system. This enables prompt customer responses even after hours, reducing the workload of employees.

[1296] Through these steps, the present invention is a system that can significantly improve the efficiency of customer service in physical stores and improve business processes.

[1297] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1298] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1299] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1300] [Fourth embodiment]

[1301] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1302] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1303] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1304] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1305] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1306] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1307] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1308] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1309] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1310] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1311] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1312] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1313] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1314] The present invention, "Dokodemo Smile," is an information analysis tool that solves communication issues that arise due to changes in the working environment, and is composed of the following elements.

[1315] System Configuration

[1316] 1. Facial Recognition Methods

[1317] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[1318] 2. Data Analysis Methods Using Generative AI Models

[1319] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[1320] 3. Voice recognition operation

[1321] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[1322] 4. Automated response measures after business hours

[1323] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[1324] 5. Data transmission and reception means

[1325] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[1326] Program processing

[1327] The program processing of this system will be specifically explained below.

[1328] 1. Server Processing

[1329] Collection of user profiles: The server periodically collects data such as each user's name, department, work history, email content, etc. in cooperation with internal and external systems. This collected data is stored in a database.

[1330] Training the generative AI model: The server trains the generative AI model using the collected data and continues to retrain the model as new data is added.

[1331] Real-time data delivery: The server sends the latest profile information to the device in real time whenever a request is made.

[1332] 2. Terminal processing

[1333] Identifying the person in front of you: The device recognizes the face of the person in front of you from the camera image through the smart glasses and queries the server. It receives information from the server and displays the person's profile.

[1334] Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[1335] Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to generate optimal responses to inquiries and respond automatically.

[1336] 3. User Processing

[1337] Profile confirmation: Users can check the profile information of the person they are meeting through the screen of their smart glasses or device, enabling smooth communication.

[1338] Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, and sending messages.

[1339] Reduced workload: Users can reduce their workload by having AI respond to inquiries on their behalf after work or when they are off-duty.

[1340] Specific examples

[1341] For example, during a meeting, the smart glasses can recognize the face of the person they are meeting and display the person's name, department, and past email content. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[1342] summary

[1343] This invention, "Dokodemo Smile," is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automatic response after work, and data transmission and reception, in order to solve communication issues caused by changes in the work environment. This will realize smooth communication and improved work efficiency.

[1344] The processing flow will be explained below.

[1345] Step 1:

[1346] Collecting user profiles

[1347] The server connects with internal and external systems and periodically collects data such as each user's name, department, work history, email content, etc. This collected data is stored centrally in the server's database.

[1348] What happens: The server calls the API to collect the latest user information, converts it into an appropriate format, and stores it in the database.

[1349] Step 2:

[1350] Training generative AI models

[1351] The server preprocesses the collected data and trains generative AI models, and also retrains existing models as new data is added.

[1352] Specific operation: The server performs preprocessing such as data cleaning and tokenization, and then trains the AI ​​model using a GPU cluster.

[1353] Step 3:

[1354] Real-time data delivery

[1355] The server sends the latest profile information in real time in response to requests from the device.

[1356] Specific operation: The device sends a request to the server, and the server retrieves the user's information from the database and returns it to the device in JSON format.

[1357] Step 4:

[1358] Identifying the person you are facing

[1359] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, queries the server, receives information from the server, and displays the other person's profile.

[1360] What it does: The device runs a facial recognition algorithm, extracts facial features, and sends them to the server, which then returns the information.

[1361] Step 5:

[1362] Voice recognition operation

[1363] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages.

[1364] What it does: The device receives voice data, converts it to text using a speech recognition engine, and executes API requests based on that information.

[1365] Step 6:

[1366] Automating business operations with AI

[1367] The server uses the generation AI even after business hours or when the server is off-line to automatically generate optimal responses to inquiries.

[1368] How it works: The server passes the query to the AI ​​model, which generates an appropriate answer and replies via email or messaging system.

[1369] Step 7:

[1370] Profile confirmation

[1371] Users can check the profile information of the person they are meeting on the smart glasses display, enabling smooth communication.

[1372] Specific operation: The user decides what to talk about and what questions to ask based on the information displayed on the smart glasses.

[1373] Step 8:

[1374] Control with voice commands

[1375] Users can issue voice commands to the device to search for employees, make calls, send messages, and more.

[1376] Specific operation: The device recognizes the voice command given by the user and performs the appropriate operation.

[1377] Step 9:

[1378] Reduced workload

[1379] Users can reduce their workload by having AI take over after work or when they are off-duty.

[1380] Specific behavior: AI automatically responds to inquiries, reducing the tasks that users have to perform.

[1381] Example 1

[1382] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1383] As the work environment changes, communication between employees is becoming less smooth. The constant wearing of masks and the introduction of a free address system can make facial recognition difficult, making face-to-face information sharing inconvenient. Furthermore, if work continues uninterrupted after the end of the working day, there is also the problem of infringing on employees' private time.

[1384] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1385] In this invention, the server includes a face recognition means, a data analysis means using a machine learning model, an operation means for analyzing voice input, an automatic response means after work hours, and an information transmission and reception means. This allows employees to easily identify the person they are meeting and efficiently operate using voice commands. In addition, since the AI ​​automatically responds after work hours, employees can protect their private time.

[1386] "Facial identification means" refers to a technical means for identifying the face of the target person using video equipment.

[1387] "Means for data analysis using machine learning models" refers to technical means for performing data analysis using machine learning algorithms based on collected data.

[1388] "Operation means for analyzing voice input" refers to a technical means for analyzing a user's voice, converting it into text, and operating the system according to the commands.

[1389] "Automatic response measures after work is completed" refers to a technical measure that allows the generation AI to automatically respond appropriately to inquiries even after the user's work is completed.

[1390] "Means for transmitting and receiving information" refers to the technical means for transmitting and receiving information between different devices and servers inside and outside the system.

[1391] "Video equipment" is a general term for hardware devices used to capture images, such as cameras and sensors.

[1392] A "machine learning algorithm" is a mathematical model that learns patterns from large amounts of data and automates tasks such as prediction and classification.

[1393] A "voice command" is an instruction to operate the system using voice.

[1394] The present invention, "Dokodemo Smile," is an information analysis system that combines multiple technical means to solve communication issues that arise with changes in the work environment. The present invention is composed of the following main components:

[1395] Facial Identification Method

[1396] The device has the ability to identify the face of the person in front of it using video equipment such as cameras and sensors. Facial identification uses OpenCV and Microsoft Azure's Face API. This method can eliminate the difficulty of face identification due to mask wearing and free address systems. For example, during a meeting, the camera in the smart glasses can detect the face of the person in front of it, convert the facial features into vectors, and send them to a server.

[1397] Data analysis methods using machine learning models

[1398] The server analyzes the collected data using machine learning algorithms (e.g., GPT-3 or BERT). This data includes the user's name, department, work history, email content, etc., and is updated regularly. The analyzed information is sent to the device in real time whenever a request is made. For example, if a user wants to check the profile of a person they are meeting with, the server will provide the latest profile information.

[1399] A means of operation that analyzes voice input

[1400] The device analyzes the user's voice, converts it into text, and operates the system according to the commands. This voice recognition function uses Google Cloud Speech-to-Text and Amazon Lex. Users can use voice commands to perform operations such as searching for employees, making calls, and sending messages. For example, if you issue the voice command "View Tanaka's profile," the voice will be converted into text and the device will perform the corresponding operation.

[1401] Automated response measures after business hours

[1402] The device uses generative AI (e.g., GPT-3 or BERT) to generate optimal responses to inquiries after work hours or while in off-mode, and automatically responds. This allows users to spend their time with peace of mind after work. For example, if a user receives an inquiry email asking, "Please tell me more about tomorrow's meeting," while in off-mode, the AI ​​will automatically generate and send a reply.

[1403] Means of sending and receiving information

[1404] The server and terminals send and receive information between different devices and servers inside and outside the system. For this purpose, a real-time data streaming platform such as Apache Kafka is used. This makes it possible to obtain the necessary information in real time and display it on the terminal. For example, if new profile data is sent from the server during a meeting, it will be displayed on the terminal immediately.

[1405] Specific examples

[1406] For example, during a meeting, a user can recognize the face of the person they are meeting and display the person's name, department, and past email content on the smart glasses. This allows the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents by email. Furthermore, the AI ​​can automatically respond to inquiries even after work hours, allowing them to make effective use of their time after work.

[1407] Examples of prompt statements

[1408] "Facial recognition is performed during meetings, and the other person's name, department, and past email content are displayed. The conversation proceeds smoothly, and documents can be emailed with subsequent voice commands. AI can automatically respond to inquiries even after work has finished."

[1409] In this way, the present invention, "Smile Anywhere," integrates a variety of technological means and responds to changes in the work environment, thereby realizing smooth communication and improved work efficiency.

[1410] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1411] Program processing flow and specific explanation

[1412] Step 1: Collecting user profiles and storing them in a database

[1413] The server periodically collects data such as user name, department, work history, email content, etc. by linking with internal and external systems. This linking is done using REST API and SOAP.

[1414] Input: User name, department, work history, email content, etc. obtained through API calls

[1415] Data processing: Analyzes acquired data in JSON or XML format and converts it into a format that can be stored in a database.

[1416] Output: Correctly formatted user data stored in the database

[1417] What it does: Every night, the server uses a job scheduler (e.g., cron) to send a request to a specified API endpoint to retrieve the latest user data, which is then parsed and stored in a database.

[1418] Step 2: Train and update the generative AI model

[1419] The server uses the collected data to train a generative AI (e.g., GPT-3, BERT) and updates the model whenever new data is added.

[1420] Input: Latest user data stored in the database

[1421] Data computation: Using machine learning algorithms to train models and optimize parameters

[1422] Output: An updated generative AI model

[1423] How it works: Every weekend, the server begins retraining the model using newly collected data. It uses a machine learning framework (e.g., TensorFlow or PyTorch) to learn patterns from large amounts of data and improve the model's accuracy.

[1424] Step 3: Identify your contact and view their profile

[1425] The device recognizes the face of the person it is meeting from the camera image via the smart glasses, queries the server and displays their profile.

[1426] Input: Smartglasses camera image

[1427] Data processing: Detecting face regions from video and generating face feature vectors

[1428] Output: Send the facial feature vector to the server, retrieve the corresponding person's profile information, and display it on the smart glasses.

[1429] How it works: During a meeting, the smart glasses analyze the camera footage in real time, use a facial recognition API to identify the face of the person they are meeting with, query the server to obtain the person's profile information, and display it on the smart glasses' display.

[1430] Step 4: Control with voice input

[1431] The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees, making calls, and sending messages.

[1432] Input: User's voice

[1433] Data processing: Converting voice to text and parsing it as commands

[1434] Output: Based on the interpreted command, perform the corresponding operation on the terminal.

[1435] What happens: When a user says "View Tanaka's profile," the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the speech to text and take the appropriate action.

[1436] Step 5: Automated responses after the business day ends

[1437] The device responds to inquiries after work hours or in off-mode using the generative AI model, generating optimal answers and responding automatically.

[1438] Input: User inquiry (e.g., email)

[1439] Data Computation: Using generative AI models to generate relevant answers

[1440] Output: Send the generated answer to the user

[1441] Specific operation: When a user receives an inquiry email after work asking, "Please tell me the details about tomorrow's meeting," the device uses the generative AI model to automatically generate and send a response saying, "Tomorrow's meeting will be held in conference room A from 10:00."

[1442] Step 6: Send and receive information in real time

[1443] The server and terminal use a real-time data streaming platform such as Apache Kafka to send and receive information in real time.

[1444] Input: Newly acquired data and updated information

[1445] Data processing: process data in real time and send it to the device in the appropriate format

[1446] Output: Real-time updated information is displayed on the terminal.

[1447] Specific operation: When new profile data is sent from the server during a meeting, it is immediately displayed on the device, allowing the user to view the latest information.

[1448] In this way, the present invention "Smile Anywhere" smoothly solves communication issues that arise with changes in the work environment through multiple processing steps.

[1449] (Application example 1)

[1450] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1451] In recent years, there has been a demand for personalized service for each customer in brick-and-mortar stores. However, it is difficult to instantly grasp a customer's name, face, and purchase history, and there is a lack of systems to ensure smooth customer service. In addition, staff are required to respond to customer inquiries even after work hours, which increases the burden on staff. There is a need for an efficient customer service support system to solve these issues.

[1452] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1453] In this invention, the server includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after work is completed, a data transmission and reception unit, a unit for acquiring customer profile information, and a unit for analyzing purchase history. This enables personalized responses to each customer in a physical store, reducing the workload of staff and improving customer satisfaction.

[1454] "Facial recognition means" is a function that uses a camera or sensor to capture and identify a person's face.

[1455] "Data analysis means using generative AI models" is a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[1456] "Operation means using voice recognition" is a function that analyzes voice commands, converts them into text, and operates the system according to those commands.

[1457] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[1458] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[1459] "Means for obtaining customer profile information" refers to a function that collects data such as customer name, affiliation, purchase history, and survey responses, and stores it in a database.

[1460] "Means for analyzing purchasing history" refers to a function that analyzes a customer's past purchasing history and understands their preferences and trends.

[1461] The following detailed description of the embodiments of the present invention will be given. Note that the embodiments described herein embody the technical features included in the claims, but do not limit the technical scope of the invention.

[1462] A customer support system in a physical store mainly consists of three components: a server, a terminal, and a user. The roles and operations of each component are as follows:

[1463] 1. Server Roles and Operations

[1464] The server plays a central role in acquiring customer profile information and analyzing the data using generative AI models. The server operates using the following hardware and software:

[1465] Hardware: Server equipment with high-performance processors, memory, and large storage capacity.

[1466] Software: Database management systems (e.g., MySQL), deep learning frameworks (e.g., TensorFlow), speech recognition APIs (e.g., Google Speech-to-Text).

[1467] The server performs the following process:

[1468] Collecting customer profiles: The server periodically collects data such as customer names, purchase history, and survey responses and stores it in a database.

[1469] Training the generative AI model: The server trains the generative AI model using the collected data and retrains the model whenever new data is added.

[1470] Real-time data delivery: The server sends the latest customer profile information to the device in real time upon request.

[1471] 2. Roles and Functions of the Device

[1472] Terminals are devices used by customer-facing staff, such as smart glasses and tablets, that operate using the following hardware and software:

[1473] Hardware: Smart glasses, camera, microphone.

[1474] Software: Facial recognition libraries (e.g., OpenCV), speech recognition software (e.g., Google Speech-to-Text).

[1475] The terminal performs the following process:

[1476] Customer face identification: The device recognizes the customer's face from the camera image of the smart glasses and queries the server for that information. The customer profile information sent from the server is displayed on the smart glasses display.

[1477] Voice recognition operation: The terminal recognizes voice commands and performs operations such as displaying customer information, checking inventory, and operating the cash register.

[1478] Automatic response after business hours: The device will use generative AI to automatically respond to inquiries after business hours or while in off mode.

[1479] 3. User Roles and Actions

[1480] The users of this system are the service staff who check customer information through the smart glasses or the terminal screen and perform operations using voice commands.

[1481] Check customer information: Staff can view information such as customer profiles and purchase history through the smart glasses screen.

[1482] Voice command operation: Staff use voice commands to search for information, check inventory, operate the cash register, and more.

[1483] Reduced workload: After the work is completed, AI automatically handles customer support, reducing the workload of staff.

[1484] Specific examples

[1485] For example, a sales staff member at a physical store can wear smart glasses and recognize the face of a customer who visits the store. The system displays the customer's profile and purchase history on the smart glasses' display, allowing the staff member to recommend products based on the customer's preferences and past purchase history. Furthermore, when the staff member issues a voice command such as "Show me recommended products for this customer," the system uses AI to select and display the appropriate products.

[1486] This will enable personalized service in physical stores, improving customer satisfaction. In addition, the AI ​​will automatically respond to customers after the store has finished its work, reducing the workload of staff.

[1487] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1488] Step 1:

[1489] The server collects customer profile information and stores it in a database. Specifically, it periodically obtains customer names, purchase histories, survey responses, etc. from the POS system in the physical store and the online survey system. The input to this process is data from the POS system and the survey system, and the output is customer information stored in the database.

[1490] Step 2:

[1491] The server uses the stored customer information to train a generative AI model. Specifically, it uses a deep learning framework (e.g., TensorFlow) to analyze customer purchasing trends and preferences. The input to this process is the customer information in the database, and the output is the trained AI model.

[1492] Step 3:

[1493] The server receives requests from the smart glasses and delivers real-time customer profile information. Specifically, it receives facial recognition data sent from the smart glasses' camera, retrieves the corresponding customer information from the database, and sends it to the smart glasses. The input of this process is facial recognition data, and the output is customer profile information.

[1494] Step 4:

[1495] The device (smart glasses) recognizes the customer's face using camera images. Specifically, it uses a facial recognition library (e.g., OpenCV) to extract facial features from the image and sends that information to a server. The input to this process is the camera image, and the output is facial recognition data.

[1496] Step 5:

[1497] The terminal recognizes the voice command and performs the required action. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text) to convert the voice command into text and follows the instructions to display customer information or check inventory. The input to this process is the voice command, and the output is the text command and the result of the action taken.

[1498] Step 6:

[1499] The terminal automatically responds to inquiries even after work hours. Specifically, it uses generative AI to generate appropriate responses to customer inquiries and automatically replies. The input to this process is the customer inquiry, and the output is the generated response.

[1500] Step 7:

[1501] The user checks customer information through the smart glasses screen and uses voice commands. Specifically, the user reads the information displayed on the smart glasses display and takes the necessary action to respond to the customer. The input of this process is the customer information displayed on the display, and the output is the user's specific action.

[1502] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1503] This invention is an information analysis tool that combines "Smile Anywhere" with an emotion engine to recognize user emotions and further improve smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[1504] System Configuration

[1505] 1. Facial Recognition Methods

[1506] This function uses cameras and sensors to capture the face of the person you are facing and identify them. By using this method, it is possible to identify the other person even in situations where face identification is difficult due to mask wearing or free address systems.

[1507] 2. Data Analysis Methods Using Generative AI Models

[1508] This function uses AI models based on machine learning and deep learning to analyze collected data and generate useful information, making it possible to present relevant information such as the other party's work history and email content.

[1509] 3. Voice recognition operation

[1510] This function analyzes the user's voice, converts it into text, and operates the system according to the commands, allowing hands-free employee search and message sending.

[1511] 4. Automated response measures after business hours

[1512] This function allows the AI ​​to automatically respond appropriately to inquiries even after the user has finished work, allowing users to use their time more freely after work.

[1513] 5. Data transmission and reception means

[1514] This is a function that sends and receives data between different devices and servers inside and outside the system, making it possible to obtain necessary information in real time and display it on the appropriate device.

[1515] 6. Emotion Engine

[1516] This function analyzes the user's voice and video data to recognize their emotional state, enabling them to respond and provide feedback according to their emotions.

[1517] Program processing

[1518] The system program of the present invention performs the following processes.

[1519] 1. Server Processing

[1520] - User profile collection: The server periodically collects data such as each user's name, department, work history, and email content in cooperation with internal and external systems. This data is then centralized in the server's database.

[1521] - Training generative AI models: The server trains generative AI models using collected data and continues to retrain existing models as new data is added.

[1522] - Training the emotion engine: The server trains the emotion engine based on the accumulated audio and video data.

[1523] - Real-time data delivery: The server sends the latest profile information and emotional state in real time in response to requests from the device.

[1524] 2. Terminal processing

[1525] - Face-to-face identification: The device recognizes the face of the person it is facing from the camera image through the smart glasses and queries the server. It receives information from the server and displays the face's profile and emotional state.

[1526] - Voice recognition operation: The device uses voice recognition to interpret the user's voice commands and perform operations such as searching for employees and sending messages. It also sends the voice data to an emotion engine for emotion recognition.

[1527] - Automated business responses using AI: The device uses generative AI after work hours or when in off-mode to automatically generate the optimal response based on the inquiry and emotion.

[1528] 3. User Processing

[1529] - Profile Check: Users can check the profile information and emotional state of the person they are meeting through the smart glasses or device screen, enabling smooth communication.

[1530] - Operate with voice commands: Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and also check the results of voice emotion analysis.

[1531] - Reduced workload: After work or while users are off-duty, AI can answer inquiries on their behalf and provide feedback based on emotion recognition, reducing their workload.

[1532] Specific examples

[1533] For example, during a meeting, a user can recognize the face of the person they are meeting and display the other person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, AI can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work.

[1534] summary

[1535] This invention is a system that integrates various methods, such as facial recognition, data analysis using generative AI models, voice recognition, automated responses after work is completed, data transmission and reception, and an emotion engine, to resolve communication barriers caused by changes in the work environment, thereby achieving smoother and more efficient communication and improved work efficiency.

[1536] The processing flow will be explained below.

[1537] Step 1:

[1538] Collecting user profiles

[1539] The server connects with internal and external systems and periodically collects each user's name, department, work history, email content, voice data, etc. This collected data is stored in a centralized database.

[1540] Specific operation: The server calls the API to collect the latest user information, converts it into an appropriate format, and saves it in the database.

[1541] Step 2:

[1542] Training generative AI models and emotion engines

[1543] The server preprocesses the collected data and trains the generative AI model and emotion engine, and retrains the AI ​​model and emotion engine every time new data is added.

[1544] How it works: The server performs preprocessing such as data cleaning and tokenization, and uses a GPU cluster to train the AI ​​model and emotion engine.

[1545] Step 3:

[1546] Real-time data delivery

[1547] The server transmits the latest profile information and emotional state in real time in response to requests from the device.

[1548] Specific operation: The device sends a request to the server, and the server retrieves the user's information and emotional state from the database and returns it to the device in JSON format.

[1549] Step 4:

[1550] Face-to-face identification and emotion recognition

[1551] The device recognizes the face of the person it is meeting from the camera image through the smart glasses, and then uses an emotion engine to recognize the other person's emotional state. It then queries the server, receives the other person's profile information and emotional state, and displays them.

[1552] How it works: The device captures the face of the person it's facing with a camera, runs a facial recognition algorithm and emotion engine, and sends the extracted facial features and emotion information to the server, which then returns the relevant information.

[1553] Step 5:

[1554] Voice recognition operation

[1555] The device uses voice recognition to analyze the user's voice commands and perform operations such as searching for employees and sending messages. The voice data is sent to an emotion engine to analyze the user's emotions.

[1556] What it does: The device receives voice data, converts it to text using a speech recognition engine, executes an API request based on that information, and analyzes the voice emotion using an emotion engine.

[1557] Step 6:

[1558] Automating business operations with AI

[1559] The server uses AI generation after work hours or when the user is off-duty to automatically generate the optimal response based on the inquiry and the user's emotional state.

[1560] How it works: The server passes the query and sentiment data to the model, generates the best answer, and sends it back via email or messaging system.

[1561] Step 7:

[1562] Profile and sentiment review

[1563] Users can check the profile information and emotional state of the person they are meeting through the screen of their smart glasses or device, allowing for smooth communication.

[1564] Specific behavior: The user sees the profile and emotional state of the other person displayed on the smart glasses and responds or engages in conversation appropriately.

[1565] Step 8:

[1566] Control and confirm emotions with voice commands

[1567] Users can issue voice commands to the device to perform operations such as searching for employees, making calls, sending messages, and checking their own and the other person's emotional state.

[1568] Specific operation: The user issues a voice command, the device recognizes it, and performs the necessary operation, displaying emotional information analyzed by the emotion engine.

[1569] Step 9:

[1570] Reduced workload

[1571] The AI ​​can take over after work or while the user is off-duty, reducing the burden on the user. Feedback based on emotion recognition also reduces stress and enables appropriate work responses.

[1572] Specific operation: AI automatically generates and sends appropriate responses to inquiries. It analyzes emotional information and provides appropriate feedback to users.

[1573] Example 2

[1574] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1575] In recent years, changes in the work environment have made it increasingly difficult for employees to communicate smoothly. Furthermore, remote work and flexible office settings have reduced face-to-face communication, making it difficult to recognize emotions and streamline work. This can lead to delays in work progress and team cooperation. A system is needed to resolve these issues, promote smooth communication between employees, and improve work efficiency.

[1576] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1577] In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, an emotion engine means, a user profile collection means, a generative AI model training means, an emotion engine training means, a real-time data distribution means, a face-to-face person identification means, and an AI-based automation means for work responses. This enables smooth communication between employees and improves work efficiency.

[1578] "Facial recognition means" is a technology that uses cameras and sensors to identify a person's face.

[1579] "Data analysis methods using generative AI models" is a technology that utilizes AI models that use machine learning and deep learning to analyze collected data and generate useful information.

[1580] "Operation means using voice recognition" is a technology that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[1581] "Automatic response measures after work is completed" refers to technology that enables the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[1582] "Data transmission / reception means" refers to technology for transmitting and receiving data between different devices and servers inside and outside the system.

[1583] The "emotion engine means" is a technology for analyzing the user's voice and video data and recognizing the user's emotional state.

[1584] "Means for collecting user profiles" refers to technology that allows the server to collect data such as the user's name, department, work history, and email content from systems both inside and outside the company.

[1585] A "generative AI model training means" is a technique for training a generative AI model using collected data and for retraining the model whenever new data is added.

[1586] "Means for training the emotion engine" refers to a technology for training the emotion engine using stored audio and video data.

[1587] "Real-time data distribution means" refers to technology for transmitting the latest user profile information and emotional state in real time in response to a request from a terminal.

[1588] The "means for identifying the person being met" is a technology that recognizes the face of the person being met from camera images via smart glasses and queries the server.

[1589] "AI-based automation of business operations" is a technology that uses generative AI after work hours or during off-duty periods to automatically generate optimal responses based on inquiries and emotions.

[1590] This invention is an information analysis tool that recognizes user emotions by combining an emotion engine with "Smile Anywhere," thereby improving smooth communication and business efficiency. The specific configuration of this system and the program processing are described below.

[1591] System Configuration

[1592] Facial Recognition Methods

[1593] The system uses cameras and sensors to capture the face of the person facing you and identify them. This method makes it possible to identify the other person even in situations where facial identification is difficult due to mask wearing or free address systems. Specifically, it utilizes face recognition technology (e.g., OpenCV) using camera images.

[1594] Data analysis methods using generative AI models

[1595] The system uses AI models (e.g., GPT-4) that use machine learning and deep learning to analyze the collected data and generate useful information, which can then present relevant information such as the other party's work history and email content.

[1596] Voice recognition operation

[1597] The system analyzes the user's voice, converts it into text, and operates the system according to the commands. This method allows hands-free employee search and message sending. Specifically, it uses a voice recognition service (e.g., Google Assistant).

[1598] Automated response measures after business hours

[1599] The system uses generative AI to automatically respond to inquiries after users have finished their work, giving users more freedom to use their time after work.

[1600] Data transmission and reception means

[1601] Send and receive data between different devices and servers inside and outside the system. This method makes it possible to obtain the necessary information in real time and display it on the appropriate device. Specifically, it uses RESTful APIs.

[1602] Emotion Engine

[1603] The system analyzes the user's voice and video data to recognize their emotional state, enabling it to respond and provide feedback according to their emotions.

[1604] How user profiles are collected

[1605] The server connects with internal and external systems to periodically collect data such as each user's name, department, work history, and email content, and centralizes it in a database. Specifically, it obtains data from LDAP directories and email servers via API.

[1606] A means of training generative AI models

[1607] The server uses the collected data to train a generative AI model, and continues to retrain the model as new data is added.

[1608] A means of training the emotion engine

[1609] The server trains the emotion engine based on the stored audio and video data. Specifically, it converts the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeds it into the emotion analysis model.

[1610] Real-time data delivery method

[1611] The server receives requests from the device and sends the latest user profile information and emotional state in real time. Specifically, it receives requests from the device using a RESTful API and returns them in JSON format.

[1612] Means of identifying the person you are facing

[1613] The device recognizes the face of the person in front of it from the camera image through the smart glasses, queries the server, and receives information from the server to display the other person's profile and emotional state. Specifically, it uses facial recognition software (e.g., OpenCV).

[1614] Automating business processes using AI

[1615] The device uses generative AI to automatically generate the optimal response based on the inquiry and the customer's emotion after work hours or when the device is in off mode. The AI ​​model is called via API, and the inquiry content and emotional data are input, and the generated response is automatically sent via email or chat.

[1616] Specific examples

[1617] For example, a user can recognize the face of the person they are meeting with during a meeting and display the person's name, department, past email content, and emotional state on the smart glasses, allowing the conversation to proceed smoothly. After the meeting, they can also use voice commands to send documents via email. Furthermore, after work, the AI ​​can respond appropriately based on the other person's emotional state, allowing them to make effective use of their time after work. An example prompt might be, "Please describe a system that allows a user to identify the face of the person they are meeting with during a meeting and display the other person's emotional state and work history."

[1618] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1619] Step 1:

[1620] Collecting user profiles

[1621] The server connects with internal and external systems (for example, LDAP directories and mail servers) to collect data such as each user's name, department, work history, email content, etc. This data is obtained via API and centralized in the server's database.

[1622] Input: Data from an LDAP directory or mail server.

[1623] Output: Centralized database.

[1624] Step 2:

[1625] Training generative AI models

[1626] The server uses the collected data to train a generative AI model (e.g., GPT-4), and retrains this model to keep it up to date whenever new data is added.

[1627] Input: Information from a centralized database.

[1628] Output: A trained generative AI model.

[1629] Step 3:

[1630] Training the Emotional Engine

[1631] The server trains the emotion engine based on the stored audio and video data, for example by converting the audio data into text using a speech recognition service (e.g., Google Speech-to-Text) and feeding the text data into the emotion analysis model.

[1632] Input: Audio data, video data.

[1633] Output: A trained emotion engine.

[1634] Step 4:

[1635] Identifying the person you are facing

[1636] The device recognizes the face of the person it is meeting from the camera image through the smart glasses. Using facial recognition software (e.g., OpenCV), it sends the recognized facial data to the server, which then returns the other person's profile and emotional state.

[1637] Input: Camera footage.

[1638] Output: The other person's profile information and emotional state.

[1639] Step 5:

[1640] Viewing a User Profile

[1641] Users can check the profile information and emotional state of the person they are meeting through the smart glasses, which allows for smooth communication. The information is displayed on the smart glasses' HUD.

[1642] Input: Profile information and emotional state received from the server.

[1643] Output: Information displayed on the smart glasses HUD.

[1644] Step 6:

[1645] Voice recognition operation

[1646] The device uses a voice recognition function to analyze the user's voice commands. For example, it uses a voice recognition service (e.g., Google Assistant) to convert the voice commands into text and analyzes the text data. As a result, it becomes possible to perform operations such as searching for employees and sending messages.

[1647] Input: The user's voice command.

[1648] Output: Actions based on the parsed voice command.

[1649] Step 7:

[1650] Real-time data delivery

[1651] The server receives requests from the device and sends the latest user profile information and emotional state in real time. It uses a RESTful API to receive requests, retrieve the latest information from the database, and return it in JSON format.

[1652] Input: A request from the terminal.

[1653] Output: Latest user profile information and emotional state.

[1654] Step 8:

[1655] Automating business operations with AI

[1656] The device uses generative AI after work hours or when in off-duty mode to automatically generate the optimal response based on the inquiry and emotion. For example, the AI ​​model can be called via API, the inquiry content and emotion data are input, and the generated response is automatically sent via email or chat.

[1657] Input: Enquiry content and emotion data.

[1658] Output: The generated answer.

[1659] Step 9:

[1660] Reduced workload

[1661] The AI ​​will answer inquiries on behalf of the user after work or while the user is off-duty. It also reduces the burden of work by providing feedback based on emotion recognition. For example, the user can request, "Please send me an email with your thoughts on today's meeting," and the AI ​​will automatically respond.

[1662] Input: User request.

[1663] Output: Auto-generated emails and feedback.

[1664] (Application example 2)

[1665] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1666] Customer service in modern brick-and-mortar stores requires quick and effective responses to individual customers, but achieving this is not easy. For example, it is difficult for store employees to check inventory while serving a customer, or to appropriately understand and respond to a customer's emotional state. Furthermore, employees are required to respond quickly to customer inquiries even after their shifts have finished, placing a heavy burden on them. These challenges hinder efficient customer service and the improvement of business processes.

[1667] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a data analysis means using a generative AI model, an operation means using voice recognition, an automatic response means after work is completed, a data transmission and reception means, a customer information display means using emotion recognition, and an inventory confirmation means using voice operation. This allows store employees to not only instantly check customer information and emotional state when meeting a customer, but also to quickly check inventory and suggest products using voice commands. Furthermore, realizing automatic response after work is completed using AI reduces the burden on employees and enables more efficient customer service.

[1668] "Facial recognition means" is a function that uses a camera or sensor to capture a person's face and identify a specific individual.

[1669] "Means for data analysis using generative AI models" refers to a function that analyzes data collected using AI models that use machine learning and deep learning, and generates useful information.

[1670] "Operation means using voice recognition" is a function that analyzes the user's voice, converts it into text, and operates the system according to the commands.

[1671] "Automatic response measures after work is completed" is a function that allows the generation AI to automatically respond appropriately to inquiries even after the user has completed work.

[1672] "Data transmission / reception means" is a function for transmitting and receiving data between different devices and servers inside and outside the system.

[1673] The "means for displaying customer information based on emotion recognition" is a function that analyzes the customer's voice and video data, recognizes their emotional state, and displays that information.

[1674] The "voice-operated inventory check means" is a function that allows you to check the inventory status using voice commands.

[1675] System Overview

[1676] The present invention is a system for improving the efficiency of customer service in brick-and-mortar stores. This system includes a facial recognition unit, a data analysis unit using a generative AI model, an operation unit using voice recognition, an automatic response unit after business hours are completed, a data transmission and reception unit, a customer information display unit using emotion recognition, and an inventory confirmation unit using voice operation.

[1677] Hardware and Software Use

[1678] The system uses the following hardware and software:

[1679] Hardware

[1680] Smart glasses (e.g., Google Glass, Vuzix Blade): Equipped with a camera and microphone to capture the face of the person you are facing.

[1681] High-performance camera: Captures video data for facial recognition.

[1682] Microphone: Captures voice commands and is used for voice recognition functions.

[1683] software

[1684] Operation OS: Android, iOS

[1685] Generative AI model: Data analysis using OpenAI, GPT-4

[1686] Emotion recognition engine: Affectiva

[1687] Speech Recognition API: Google Speech-to-Text

[1688] Cloud databases: Firebase, AWS RDS

[1689] Real-time data transfer: WebSocket, Firebase Realtime Database

[1690] Program processing

[1691] Acquiring customer information through facial recognition

[1692] The camera in the smart glasses captures the customer's face. The captured image data is analyzed by facial recognition and compared with a database to obtain the customer's profile information. This allows the necessary information to be displayed immediately when interacting with the customer.

[1693] Emotion-aware response

[1694] The captured video data is sent to an emotion recognition engine (Affectiva) to analyze the customer's emotional state. The resulting emotional state (e.g., satisfaction, confusion, interest, etc.) is displayed on the smart glasses' display, allowing store staff to quickly respond according to the customer's emotions.

[1695] Voice recognition operation

[1696] When an employee issues a voice command, the system uses the Google Speech-to-Text API to convert the voice into text, analyzes it, and performs the command accordingly. For example, if you issue the voice command "Check inventory, product number 12345," inventory information will be retrieved from a cloud database and displayed in real time.

[1697] Automatic response after business hours

[1698] Even after work hours, automated responses are provided using generative AI models (OpenAI, GPT-4). The AI ​​model can generate and send appropriate responses to customer inquiries, reducing the workload of employees.

[1699] Examples and prompts

[1700] Example 1: Checking inventory

[1701] When a store clerk issues a voice command through the smart glasses, such as "Check stock, product number 12345," the system retrieves the product's stock status in real time from the cloud database and displays it on the smart glasses' display.

[1702] Example 2: Customer service using sentiment analysis

[1703] An emotion recognition engine analyzes the video data captured by the smart glasses' camera, and if it indicates that the customer is confused, employees can immediately change their response.

[1704] ChatGPT prompt example

[1705] Prompt for generative AI models (OpenAI, GPT-4):

[1706] "The customer is confused. What should we do next?"

[1707] Customer profile information:

[1708] Name: Taro

[1709] Past purchase history: Many home appliances

[1710] Preferences: Products with the latest technology

[1711] Emotional state: Confused

[1712] Generation example:

[1713] "Taro, what kind of home appliances would you like to buy today? We also have the latest smart home devices."

[1714] In this way, the operation of the system can significantly improve the quality of customer service and business efficiency.

[1715] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1716] Step 1:

[1717] A user puts on the smart glasses and launches the application. The smart glasses' camera captures the customer's face, and the video data is input into the facial recognition means. The facial recognition model analyzes the video data and extracts facial feature points. The extracted feature points are sent to the server for matching with a database. The server searches the database, obtains the corresponding customer's profile information, and sends it to the terminal. The terminal displays this information on the smart glasses' display. This allows employees to check information such as the customer's name and purchase history in real time.

[1718] Step 2:

[1719] The customer's video data captured by the smart glasses is input into an emotion recognition engine (Affectiva). The emotion recognition engine analyzes the customer's facial expressions and identifies their emotional state. The identified emotional information (e.g., satisfaction, confusion, interest, etc.) is sent to the terminal and displayed on the smart glasses' display. This allows employees to quickly respond according to the customer's emotions.

[1720] Step 3:

[1721] When a user issues a voice command, the microphone in the smart glasses captures the voice. The captured voice data is input into a speech recognition API (Google Speech-to-Text) and converted into text. The converted text data is sent to the server, where it is analyzed by a generative AI model (OpenAI, GPT-4). For example, in response to the command "Check stock, product number 12345," the server searches a cloud database and retrieves stock information for product number 12345. This information is sent to the device and displayed on the smart glasses' display. This allows employees to respond quickly to customer questions without taking their eyes off the customer.

[1722] Step 4:

[1723] Even after the user has finished their work, the generative AI model (OpenAI, GPT-4) remains in standby mode on the server. When a customer makes an inquiry, the inquiry is entered into the server. The server uses the generative AI model to generate the optimal answer and sends it to the customer via an automated response system. This enables prompt customer responses even after hours, reducing the workload of employees.

[1724] Through these steps, the present invention is a system that can significantly improve the efficiency of customer service in physical stores and improve business processes.

[1725] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1726] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1727] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1728] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1729] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1730] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1731] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1732] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1733] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1734] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1735] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1736] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1737] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1738] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1739] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1740] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1741] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1742] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1743] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1744] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1745] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1746] The following is further disclosed regarding the above embodiment.

[1747] (Claim 1)

[1748] A facial recognition means;

[1749] A means of data analysis using generative AI models;

[1750] A voice recognition operation means;

[1751] Automatic response measures after business hours are over,

[1752] data transmission and reception means;

[1753] A system including:

[1754] (Claim 2)

[1755] The system according to claim 1, wherein the face of the person being confronted is recognized using camera images.

[1756] (Claim 3)

[1757] 10. The system of claim 1, wherein operations such as searching for employees and sending messages are performed in accordance with voice commands.

[1758] "Example 1"

[1759] (Claim 1)

[1760] A face identification means;

[1761] Data analysis methods using machine learning models,

[1762] an operating means for analyzing a voice input;

[1763] Automatic response measures after business hours are over,

[1764] A means for transmitting and receiving information;

[1765] A system including:

[1766] (Claim 2)

[1767] The system according to claim 1, wherein the face of the person being confronted is identified using a video device.

[1768] (Claim 3)

[1769] 2. The system according to claim 1, wherein operations such as person search and message sending are performed in accordance with voice input.

[1770] "Application Example 1"

[1771] (Claim 1)

[1772] A facial recognition means;

[1773] A means of data analysis using generative AI models;

[1774] A voice recognition operation means;

[1775] Automatic response measures after business hours are over,

[1776] data transmission and reception means;

[1777] a means for obtaining customer profile information;

[1778] A means for analyzing purchase history;

[1779] A system including:

[1780] (Claim 2)

[1781] The system according to claim 1, wherein the face of the person being confronted is recognized using camera images.

[1782] (Claim 3)

[1783] 10. The system according to claim 1, wherein information search and operation are performed according to voice commands.

[1784] "Example 2: Combining Emotion Engines"

[1785] (Claim 1)

[1786] A facial recognition means;

[1787] A means of data analysis using generative AI models;

[1788] A voice recognition operation means;

[1789] Automatic response measures after business hours are over,

[1790] data transmission and reception means;

[1791] an emotion engine means;

[1792] a means for collecting user profiles;

[1793] A means of training the generative AI model;

[1794] A means of training the emotion engine;

[1795] a means for delivering real-time data;

[1796] A means for identifying the person being confronted;

[1797] AI-based automation of business operations,

[1798] A system including:

[1799] (Claim 2)

[1800] The system according to claim 1, wherein the system uses camera footage to recognize the face of the person in front of the user and displays their profile and emotional state.

[1801] (Claim 3)

[1802] The system according to claim 1, wherein data transmission / reception and business correspondence operations are performed in accordance with voice commands.

[1803] "Application example 2 when combining emotion engines"

[1804] (Claim 1)

[1805] A facial recognition means;

[1806] A means of data analysis using generative AI models;

[1807] A voice recognition operation means;

[1808] Automatic response measures after business hours are over,

[1809] data transmission and reception means;

[1810] A means for displaying customer information using emotion recognition;

[1811] A means of checking inventory by voice operation,

[1812] A system including:

[1813] (Claim 2)

[1814] 2. The system according to claim 1, further comprising means for recognizing the face of the person being met using camera images and displaying customer information.

[1815] (Claim 3)

[1816] The system according to claim 1, further comprising an operating means for checking stock availability and searching for suggested products in accordance with voice commands. [Explanation of symbols]

[1817] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A facial recognition means; A means of data analysis using generative AI models; A voice recognition operation means; Automatic response measures after the business is over, data transmission and reception means; A system including:

2. The system according to claim 1, wherein the face of the person being confronted is recognized by using a camera image.

3. The system of claim 1, wherein operations such as searching for employees and sending messages are performed in accordance with voice commands.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A