system
A system using image and voice recognition with a large-scale language model enhances communication by suggesting optimal topics and behaviors in real-time, addressing challenges in dating and sales interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Individuals face challenges in thinking of appropriate topics and behaviors during social interactions like dating or sales calls, often leading to ineffective communication due to nervousness or lack of emotional understanding, and existing systems fail to provide real-time emotional analysis and advice.
A system utilizing image and voice recognition, combined with a large-scale language model, to estimate emotions and suggest optimal topics and behaviors through an augmented reality device, with initial setup, real-time support, and follow-up actions.
Enhances communication quality by providing real-time advice on topics and actions, improving user performance in dating and sales situations.
Smart Images

Figure 2026036235000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When dating or making a sales call, it can be difficult to think of appropriate topics to talk about or how to behave, making it difficult to communicate effectively with the other person. Furthermore, nervousness can cause people to go blank and be unable to come up with a topic to talk about. To improve these situations and provide users with a better dating or sales experience, a system is needed that analyzes emotions and situations in real time and provides optimal advice. [Means for solving the problem]
[0005] The present invention provides a system that includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, and means for displaying information on an augmented reality (AR) device worn by the user. The system also includes means for the user to register basic information about the other party through natural language input, perform initial setup in conjunction with past conversation history data, and means for proposing date plans and sales strategies. This system allows users to improve the quality of their communication by suggesting optimal topics and actions in real time. It also provides means for receiving follow-up actions after a date or sales call ends, and means for generating LINE and email messages, providing ongoing support and improving results.
[0006] "Image recognition" is a technology that analyzes and identifies specific information and patterns from acquired image data.
[0007] "Emotion estimation" is a technology that estimates a user's emotional state (joy, sadness, anger, etc.) from data such as image recognition and voice recognition.
[0008] "Speech recognition" is a technology that analyzes acquired voice data and identifies the language, emotion, and speaker.
[0009] A "large-scale language model" is a natural language processing model built on large amounts of text data, and is capable of performing tasks such as text generation, question answering, and translation.
[0010] "Topic generation" is the process of suggesting appropriate conversation topics based on specific contexts and conditions.
[0011] "Deportment" refers to the attitude and behavior in a particular situation, and is the appropriate behavior and way of expressing oneself in interpersonal relationships.
[0012] An "augmented reality (AR) device" is a device that displays virtual information overlaid on the user's real field of vision, providing additional information to the user's visual information.
[0013] "Natural language input" is a method in which a user inputs settings and instructions in natural language using a keyboard or voice input.
[0014] "Past conversation history data" is a record of conversations that have taken place between the user and the other party, and is used as reference information for the next conversation.
[0015] A "date plan" is a specific plan that includes how the date will proceed, the places to visit, and the activities to do.
[0016] A "sales strategy" refers to the specific plans and methods for successful sales activities, and how to propose products and services.
[0017] "Follow-up" refers to additional communication or actions that take place after a date or business transaction has concluded, with the aim of building and maintaining an ongoing relationship.
[0018] "Message generation for LINE and email" is a process that automatically creates content to be sent via messaging apps or email, and is used to reduce the effort required by users. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention is a system for facilitating communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. The following describes how the system of the present invention is specifically implemented.
[0041] Overall system configuration
[0042] The system consists of the following main components:
[0043] 1. Client device (augmented reality (AR) device worn by the user)
[0044] 2. Server (responsible for major processing on the cloud)
[0045] 3. User interface (smartphone app or web interface)
[0046] Initial setup and preparation phase
[0047] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[0048] Real-time support during business hours
[0049] The user puts on AR glasses or AR contact lenses, and the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. The server integrates the emotion estimation results with the current situation and generates optimal advice using a large-scale language model. This allows the user to receive suggestions for appropriate topics and actions in real time. The advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[0050] Follow-up
[0051] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user.
[0052] Specific examples
[0053] In the case of a date
[0054] 1. The day before the date
[0055] The user enters the other person's basic information into a smartphone app.
[0056] The server suggests a plan of "lunch at Cafe A" and "watch movie B."
[0057] The user selects "Watch Movie B" and the server reserves the tickets.
[0058] 2. The day of the date
[0059] The user puts on the AR glasses.
[0060] During the date, the AR glasses transmit video and audio to a server.
[0061] The server infers that the other person is "interested" based on their facial expression and tone of voice.
[0062] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[0063] The advice is displayed in the user's field of view, and the user provides the topic.
[0064] 3. After the date
[0065] The server generates a message saying, "Thank you for today. It was fun!"
[0066] The message is notified to the smartphone app, and the user confirms it before sending it.
[0067] In the case of sales
[0068] 1. Before the sales visit
[0069] The user enters the other company's information and the name of the person in charge into the smartphone app.
[0070] The server prepares initial proposals and topics based on past transaction history and related information.
[0071] 2. During a sales visit
[0072] The user puts on the AR contact lenses.
[0073] During business hours, the AR contact lenses capture video and audio and send them to a server.
[0074] The server analyzes the facial expressions and tone of voice of the person in charge and determines that they are "interested."
[0075] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[0076] The advice is displayed in the user's field of view and the user demonstrates it.
[0077] 3. After the visit
[0078] The server generates a follow-up email.
[0079] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[0080] Thus, the present invention is a system that supports users in real time in choosing the best topics and actions to take during dates or business trips, thereby improving the quality of communication.
[0081] The processing flow will be explained below.
[0082] Initial setup and preparation phase
[0083] Step 1:
[0084] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[0085] Step 2:
[0086] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received data, and retrieves matching data.
[0087] Step 3:
[0088] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[0089] Step 4:
[0090] The user checks the proposed date plans and sales strategies, selects one, and sends the selection information to the server.
[0091] Step 5:
[0092] The server makes the necessary reservations (restaurants, cinemas, etc.) based on the user's selection, and notifies the user's app that the reservation has been completed.
[0093] Real-time support during business hours
[0094] Step 1:
[0095] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[0096] Step 2:
[0097] The device uses its sensors to capture video and audio in real time, and sends the captured data to a server.
[0098] Step 3:
[0099] The server analyzes the received video data using an image recognition algorithm to infer emotions from facial expressions, and analyzes the audio data using a voice recognition algorithm to infer emotions from the tone and patterns of the voice.
[0100] Step 4:
[0101] The server sends the emotion estimation results to the user's smartphone app, which simultaneously analyzes the emotion information and the current dating or sales context using a large-scale language model to generate optimal advice.
[0102] Step 5:
[0103] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[0104] Follow-up
[0105] Step 1:
[0106] The user enters into the app that a date or business has ended, and the device sends all data up to the end to the server.
[0107] Step 2:
[0108] The server analyzes the data at the end, generates the next follow-up action to be taken, and notifies the smartphone app with the generated recommendation message.
[0109] Step 3:
[0110] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[0111] Example 1
[0112] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0113] In modern society, it is difficult to grasp appropriate topics and actions in real time in communication situations such as dating or business. In particular, accurately understanding one's own emotions and the emotions of others and acting accordingly is extremely important, but not easy to achieve. Conventional systems lack the ability to assess the other person's emotions and provide appropriate advice, resulting in a decline in the quality of communication. New technology is needed to solve these problems.
[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0115] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for acquiring video and audio in real time and transmitting them to the server, and means for analyzing all data up to the end and generating the next action. This allows the user to receive suggestions for appropriate topics and actions in real time, enabling high-quality communication even during dates or business trips.
[0116] "Image recognition" is a technology that analyzes image data acquired using cameras and sensors to identify objects, people, features, etc.
[0117] "Emotion estimation" is a technology that determines the other person's psychological state and emotions by analyzing image and audio data.
[0118] "Speech recognition" is a technology that analyzes voice data acquired using devices such as microphones and identifies what a person is saying and the characteristics of their voice.
[0119] A "large-scale language model" is an artificial intelligence model that learns from large amounts of text data to generate and understand natural language.
[0120] An "augmented reality (AR) device" is a device that displays digital information superimposed on real-world visual information.
[0121] "Real-time" means processing acquired data immediately without delay and providing results instantly.
[0122] "Capturing video and audio" means collecting visual and audio information from the surroundings using devices such as cameras and microphones.
[0123] A "server" is a central computer system that provides services over a network.
[0124] "Analyzing all data" means extracting and analyzing information for a specific purpose from all acquired data.
[0125] "Generating the next action" means proposing the next action or response that the user should take based on the analysis results.
[0126] This invention is a system for facilitating smooth communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. Below, we will explain how to specifically implement the system of this invention.
[0127] Hardware and software used
[0128] Client device: An augmented reality (AR) device worn by a user (e.g., AR glasses, AR contact lenses).
[0129] Server: A computer system that handles major processing on the cloud.
[0130] User interface: Smartphone app and web interface.
[0131] Initial setup and preparation phase
[0132] 1. Enter information
[0133] The user starts the smartphone app and enters the other person's basic information (name, age, hobbies, previous conversations), which is then received by the server and stored in a database.
[0134] 2. Information Reception and Analysis
[0135] The server analyzes the received information and generates initial topics and suggestions based on past conversation history data. The server then uses a generative AI model to propose date plans and sales strategies.
[0136] 3. Plan proposal
[0137] The server proposes specific plans to the user, such as "lunch at Cafe A" or "watch movie B." The user selects one of the proposed plans and sends it to the server, which then arranges the reservations required for the selected plan.
[0138] Example prompt sentence:
[0139] Based on the other person's basic information, propose a date plan (e.g., lunch at Cafe A, watching Movie B).
[0140] Real-time support during business hours
[0141] 1. Wearing the device
[0142] Users wear AR glasses or AR contact lenses, and these devices capture video and audio in real time.
[0143] 2. Data Acquisition and Transmission
[0144] The device sends the captured video and audio to a server, which analyzes them and uses image and voice recognition technology to estimate the other person's emotions.
[0145] 3. Emotion Estimation and Advice Generation
[0146] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[0147] 4. Displaying Advice
[0148] The server sends the generated advice to the user's device, which then displays it in the user's field of view. For example, if the person speaking is smiling, the advice displayed might be, "Try talking about a movie you recently saw."
[0149] Example prompt sentence:
[0150] Based on the video and audio captured in real time, estimate the other person's emotions (e.g., interest) and generate appropriate topics of conversation.
[0151] Follow-up
[0152] 1. Notice of Termination
[0153] The user notifies the smartphone app that a date or business meeting has ended. The server collects all data up to that point and begins analyzing it.
[0154] 2. Suggested next actions
[0155] The server generates the next action to be taken. For example, the server generates a follow-up message such as "Thank you for today. I had a great time!" and notifies the user's smartphone app. The follow-up message is sent when the user confirms and sends it.
[0156] Example prompt sentence:
[0157] Generate a follow-up message to send after the date ends (e.g., "Thank you for today. I had a great time!").
[0158] In this way, the system of the present invention can support users in real time to choose the most appropriate topic and action in any situation, thereby realizing high-quality communication even during dates or business meetings.
[0159] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0160] Step 1:
[0161] Enter information
[0162] The user launches the smartphone app and enters the other person's basic information (name, age, hobbies, and previous conversation details).
[0163] Input: The user enters the following information into the input form: "Name = Tanaka Taro, Age = 28, Hobby = Reading, Previous conversation content = We talked about movies."
[0164] Output: The data entered by the user is saved in the app and sent to the server.
[0165] Step 2:
[0166] Information reception and analysis
[0167] The server receives the basic information sent by the user.
[0168] The server analyzes the received data and stores it in a database. By linking it with past conversation history data, it generates optimal initial topics and suggestions.
[0169] Input: Basic information data submitted by the user.
[0170] Data processing: Obtain past conversation data from the database and analyze the other person's interests.
[0171] Output: Generate a list of initial topics and suggestions.
[0172] Step 3:
[0173] Plan proposal
[0174] The server uses the generated AI model to propose specific date plans and sales strategies.
[0175] Example: Generate a plan that includes "Lunch at Cafe X" and "Watch Movie Y."
[0176] Input: A candidate list and a prompt sentence for the generative AI model.
[0177] Data processing: The AI model generates a date plan based on the prompt.
[0178] Output: A list of specific plans proposed to the user.
[0179] Step 4:
[0180] Plan selection and reservation arrangements
[0181] The user selects a plan suggested by the smartphone app. For example, they select "Watch Movie Y."
[0182] The server then arranges the necessary reservations for the selected plan, in this case booking movie tickets online.
[0183] Input: The user's selected plan.
[0184] Data processing: Connect to the online reservation system based on the selected plan and complete the reservation.
[0185] Output: Reservation confirmation information is sent from the server and notified to the user.
[0186] Step 5:
[0187] Wearing the device
[0188] The user puts on AR glasses or AR contact lenses before starting a date or business trip.
[0189] Input: The act of a user putting on a device.
[0190] Output: The AR device is initialized and begins capturing video and audio in real time.
[0191] Step 6:
[0192] Acquiring and Sending Data
[0193] The device captures video and audio in real time and transmits them to the server.
[0194] Input: Surrounding video and audio information.
[0195] Data processing: The device acquires the information, encodes it, and uploads it to the server.
[0196] Output: The data stream that the server receives in real time.
[0197] Step 7:
[0198] Emotion estimation and advice generation
[0199] The server analyzes the received video and audio and uses image and voice recognition technology to estimate the other person's emotions. For example, it detects the other person's smile and infers that they are "interested."
[0200] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[0201] Input: Received video and audio data.
[0202] Data processing: Analyzes video and audio to recognize the other person's emotions. A large-scale language model generates dialogue advice based on the recognition results.
[0203] Output: The generated advice is displayed in the user's field of view.
[0204] Step 8:
[0205] Displaying Advice
[0206] The server transmits the generated advice to the user's terminal.
[0207] The device displays advice in the user's field of view, such as "Tell me about a movie you recently saw."
[0208] Input: The generated advice.
[0209] Data processing: Converting advice into a visually displayable format.
[0210] Output: An advisory message that is displayed in the user's field of view.
[0211] Step 9:
[0212] Notice of Termination
[0213] The smartphone app notifies the user when a date or business meeting has ended.
[0214] Input: User clicks the "Exit" button.
[0215] Output: A completion notice is sent to the server.
[0216] Step 10:
[0217] Data analysis
[0218] The server collects all data up to the end point and begins analyzing it, for example, analyzing the content of the conversation that day and the changes in the other person's emotions.
[0219] Input: Termination notice and all collected data.
[0220] Data processing: Analyze conversation content and emotional data to extract trends.
[0221] Output: A report of the analysis results.
[0222] Step 11:
[0223] Suggested next actions
[0224] The server generates the next action to be taken, for example, a follow-up message such as "Thank you for today. I had a great time!"
[0225] Input: Analysis results.
[0226] Data processing: The next action is generated using a generative AI model.
[0227] Output: The actions and messages generated.
[0228] Step 12:
[0229] Notification and confirmation
[0230] The server generates a message and sends it to the user's smartphone app. The user confirms it and sends the message.
[0231] Input: The generated message.
[0232] Data processing: Notify the message and ask for confirmation.
[0233] Output: The message that is confirmed by the user and sent to the other party.
[0234] (Application example 1)
[0235] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0236] Existing customer service support systems make it difficult for staff to provide optimal product recommendations and customer service methods for each individual customer in a physical store. This makes it difficult to provide effective customer service to increase customer satisfaction, and it is difficult to increase store sales. In addition, because the system relies on the customer service skills of the staff, it is difficult for new or inexperienced staff to provide effective customer service.
[0237] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0238] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for analyzing the customer's facial expressions and acquiring facial emotion data, means for analyzing voice data and converting it into text, means for generating appropriate product proposals by combining the text data and the facial emotion data, and means for synthesizing the generated product proposals and notifying the staff. This enables the staff to grasp the customer's emotions and conversation content in real time and immediately receive optimal product proposals and customer service methods.
[0239] "Image recognition" is a technology that detects and analyzes the characteristics of specific objects or people from image data.
[0240] "Emotion estimation" means estimating a person's emotional state (e.g., joy, anger, sadness, etc.) based on image and audio data.
[0241] "Speech recognition" is a technology that analyzes voice data and converts it into text data.
[0242] A "large-scale language model" is an artificial intelligence model trained on vast amounts of text data that is capable of understanding and generating natural language.
[0243] An "augmented reality (AR) device" is a device that displays computer graphics overlaid on images of the real world.
[0244] "Analyzing facial expressions to obtain facial emotion data" means using image recognition technology to extract emotional information from a person's facial expressions and obtain that data.
[0245] "Analyzing voice data and converting it into text" means converting voice data into text data using voice recognition technology.
[0246] "Generating appropriate product proposals by combining text data and facial emotion data" means automatically generating optimal product proposals for customers using text data and facial emotion data.
[0247] "Synthesize voice and notify staff" means synthesizing the generated product suggestions and advice into voice and conveying it to staff.
[0248] This invention provides a system that makes full use of image recognition, voice recognition, and large-scale language models to support customer service in brick-and-mortar stores. Specific embodiments for implementing this system are described below.
[0249] System configuration
[0250] The system consists of the following main components:
[0251] 1. Client devices (including smart glasses and head-mounted displays)
[0252] 2. Server (located on the cloud and responsible for main processing)
[0253] 3. User interface (smartphone app or web interface)
[0254] Program Overview
[0255] Emotion estimation using image recognition
[0256] The client device captures the customer's facial expressions in real time while serving customers in a physical store and sends the image data to a server. The server then uses image recognition technology such as Amazon Rekognition to analyze the customer's emotions (e.g., joy, anger, sadness, etc.) and obtains facial emotion data.
[0257] Emotion estimation using speech recognition
[0258] The client terminal records the conversation with the customer and sends the audio data to the server. The server then converts the audio data into text data using voice recognition technology such as the Google® Speech-to-Text API. This text data is then used to estimate emotions based on the conversation.
[0259] Generating optimal suggestions using large-scale language models
[0260] The server uses a large-scale language model (e.g., GPT-4 (registered trademark)) to generate optimal product suggestions and customer service methods based on text data and facial emotion data. The generated suggestions are input as prompt sentences as follows:
[0261] Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation: "What products are you interested in recently?"
[0262] Display suggestions and speech synthesis
[0263] The generated suggestions are notified to staff using voice synthesis technology (such as pyttsx3). The user (staff member) receives this information in real time through smart glasses or a head-mounted display attached to the client device and can take appropriate action.
[0264] Specific examples
[0265] Example 1: Proposing a new product
[0266] When a customer asks a staff member, "What products do you recommend?", the system converts the voice data into text data and analyzes facial expressions to infer that the customer is "interested." Based on this information, a large-scale language model generates a suggestion such as "Recommend a newly arrived product," and notifies the staff member via voice synthesis.
[0267] Example 2: Follow-up suggestions
[0268] After the staff member has finished serving the customer, the system analyzes the acquired data and generates an appropriate follow-up action (e.g., sending a thank-you message) and notifies the user interface. The user can then confirm the suggestion and take action if necessary.
[0269] As a result, this system allows staff to understand customers' emotions and conversation content in real time, enabling them to instantly receive optimal product suggestions and customer service methods.
[0270] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0271] Step 1:
[0272] The user wears the client terminal.
[0273] How it works: The customer wears smart glasses or a head-mounted display, which allows the device's camera and microphone to capture the customer's video and audio in real time.
[0274] Input: The client device worn by the user.
[0275] Output: Real-time video and audio capture.
[0276] Step 2:
[0277] The terminal captures the customer's video and audio and sends them to the server.
[0278] How it works: The device's camera captures the customer's facial expressions, and the microphone records the conversation. This data is then sent to a cloud server.
[0279] Input: Video and audio data acquired by the device.
[0280] Output: Video and audio data sent to the server.
[0281] Step 3:
[0282] The server uses image recognition technology to analyze the customer's facial expressions and obtain facial emotion data.
[0283] How it works: The server runs an image recognition algorithm such as Amazon Rekognition to estimate the customer's emotions from their facial expressions. For example, the results can be "happy" or "interested" based on the facial expression.
[0284] Input: Video data sent to the server.
[0285] Output: Facial emotion data (e.g., "HAPPY", confidence level 95.0%).
[0286] Step 4:
[0287] The server uses voice recognition technology to convert the voice data into text data.
[0288] What happens: The server runs a speech recognition algorithm, such as the Google Speech-to-Text API, to convert the conversation into text. For example, it might generate something like, "What products are you interested in these days?"
[0289] Input: The audio data sent to the server.
[0290] Output: Text data (e.g., "What products have you been interested in recently?").
[0291] Step 5:
[0292] The server combines text data and facial emotion data to input prompts into a large-scale language model to generate optimal product suggestions.
[0293] Specific operation: The server uses a large-scale language model (e.g., GPT-4) to generate a prompt sentence and then generates appropriate product suggestions based on it. An example of a prompt sentence is "Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation content: 'What products are you interested in recently?'"
[0294] Input: Text data and facial emotion data.
[0295] Output: Product suggestions (e.g. "Recommend new products").
[0296] Step 6:
[0297] The product suggestions generated by the server are notified to staff using voice synthesis technology.
[0298] Specific operation: The server uses a speech synthesis algorithm such as pyttsx3 to convert the generated product suggestions into voice and send it to the client terminal. The user (staff member) receives the suggestions by voice and can respond appropriately.
[0299] Input: Server-generated product suggestions.
[0300] Output: A voice-synthesized product suggestion notification.
[0301] Step 7:
[0302] Based on the user-generated proposals, appropriate products are proposed to the customer.
[0303] Specific operation: The user (staff member) uses the product proposal notified by voice from the client terminal as a reference to explain and propose new products to the customer.
[0304] Input: Speech-synthesized product suggestions.
[0305] Output: Product recommendations to customers.
[0306] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0307] This invention is a system that analyzes the emotions of users and their partners in real time during dates and business meetings, and suggests optimal topics and behaviors. This system achieves highly accurate emotion analysis by combining image recognition, speech recognition technology, large-scale language models, and an emotion engine.
[0308] Overall system configuration
[0309] The system consists of the following main components:
[0310] 1. Client device (augmented reality (AR) device worn by the user)
[0311] 2. Server (responsible for major processing on the cloud)
[0312] 3. Emotion engine (analyzing user and other party emotions in real time)
[0313] 4. User interface (smartphone app or web interface)
[0314] Initial setup and preparation phase
[0315] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[0316] Real-time support during business hours
[0317] When a user wears AR glasses or AR contact lenses, the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. At the same time, an emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[0318] The server integrates the emotional information of both the user and the other party, analyzes the current dating or sales context using a large-scale language model, and generates optimal advice. The generated advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[0319] Furthermore, the emotion engine learns from the user's past emotional data, analyzes their emotional tendencies, and provides advice in response to predicted emotional changes. This makes it easier for the user to understand their own emotional state and continue the conversation calmly.
[0320] Follow-up
[0321] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user. The notified message can be sent after the user confirms it.
[0322] Specific examples
[0323] In the case of a date
[0324] 1. The day before the date
[0325] The user enters the other person's basic information into a smartphone app.
[0326] The server suggests plans such as "lunch at Cafe A" or "watching Movie B."
[0327] The user selects "Watch Movie B" and the server reserves the tickets.
[0328] 2. The day of the date
[0329] The user puts on the AR glasses.
[0330] During the date, the AR glasses transmit video and audio to a server.
[0331] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[0332] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[0333] The advice is displayed in the user's field of view, and the user provides the topic.
[0334] 3. After the date
[0335] The server generates a message saying, "Thank you for today. It was fun!"
[0336] The message is notified to the smartphone app, and the user confirms it before sending it.
[0337] In the case of sales
[0338] 1. Before the sales visit
[0339] The user enters the other company's information and the name of the person in charge into the smartphone app.
[0340] The server prepares initial proposals and topics based on past transaction history and related information.
[0341] 2. During a sales visit
[0342] The user puts on the AR contact lenses.
[0343] During business hours, the AR contact lenses capture video and audio and send them to a server.
[0344] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[0345] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[0346] The advice is displayed in the user's field of view and the user demonstrates it.
[0347] 3. After the visit
[0348] The server generates a follow-up email.
[0349] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[0350] In this way, the present invention is a system that improves the quality of communication by analyzing the emotions of the user and the other party in real time and suggesting optimal topics and actions. The introduction of an emotion engine makes it easier to understand the user's own emotions, enabling more effective support.
[0351] The processing flow will be explained below.
[0352] Initial setup and preparation phase
[0353] Step 1:
[0354] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[0355] Step 2:
[0356] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received information, and retrieves matching data.
[0357] Step 3:
[0358] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[0359] Step 4:
[0360] The user checks the proposed date plans and sales strategies and selects one. The selected information is sent to the server.
[0361] Step 5:
[0362] The server makes the necessary reservations (for restaurants, cinemas, etc.) based on the user's selection, and notifies the user's smartphone app that the reservation has been completed.
[0363] Real-time support during business hours
[0364] Step 1:
[0365] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[0366] Step 2:
[0367] The device uses its sensors to capture video and audio in real time, and the captured data is sent to a server.
[0368] Step 3:
[0369] The server analyzes the video data it receives using an image recognition algorithm to infer emotions from facial expressions, and it also analyzes audio data using a voice recognition algorithm to infer emotions from the tone and patterns of voice.
[0370] Step 4:
[0371] The server creates the other person's emotional information based on the emotion estimation results. At the same time, the emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[0372] Step 5:
[0373] The server integrates the emotional information of both the user and the other party and generates optimal advice using a large-scale language model. The advice is then displayed in the user's field of view.
[0374] Step 6:
[0375] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[0376] Follow-up
[0377] Step 1:
[0378] The user enters into the app that a date or business trip has ended, and the device sends all data up to the end to the server.
[0379] Step 2:
[0380] The server analyzes the data at the end and generates a recommended action to be taken next, which is then sent to the user's smartphone app.
[0381] Step 3:
[0382] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[0383] Example 2
[0384] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0385] In conventional dating and sales situations, it has been difficult for users to accurately read the other person's emotions and decide on appropriate topics and actions. Furthermore, it has been difficult for users to manage their own emotions while having an effective conversation. This has led to a decline in the quality of communication and a decrease in the success rate of dates and sales.
[0386] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through voice recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and an emotion engine for analyzing the user's emotions in real time. This configuration enables the user to grasp the emotions of others and themselves in real time and take optimal actions and responses.
[0387] "Image recognition" is the technology that extracts information from digital images and videos and identifies specific patterns and features.
[0388] "Speech recognition" is a technology that analyzes voice data and converts spoken words into text data.
[0389] A "large-scale language model" is an artificial intelligence model trained on large amounts of text data, and is a model that excels at understanding and generating natural language.
[0390] An "augmented reality (AR) device" is a device that overlays digital information onto a user's real-world field of vision. Examples include AR glasses and AR contact lenses.
[0391] The "emotion engine" is a system that analyzes the emotional state of the user and the other person in real time and provides appropriate feedback and advice based on the results.
[0392] A "date plan" refers to a proposal or action plan for systematically preparing the progress of a date.
[0393] A "sales strategy" is a plan for effective contact methods and presentations in sales activities.
[0394] "Real-time data" is data that represents ongoing events or conditions and should be processed immediately.
[0395] "Follow-up actions" are subsequent actions that should be taken after a date or business meeting, with the purpose of maintaining and strengthening the relationship with the other person.
[0396] "Initial settings" refers to the setup work required before starting to use the system, in which basic information about the user and the other party is entered and preparations required for system operation are made.
[0397] The system based on this invention analyzes the emotions of the user and the other party in real time during dating or business, and suggests optimal topics and behaviors. This system includes the following main components: image recognition, speech recognition, a large-scale language model, an augmented reality (AR) device, and an emotion engine.
[0398] Hardware and software used
[0399] Hardware: Augmented reality (AR) glasses, augmented reality (AR) contact lenses, smartphones
[0400] Software: Image recognition software (e.g., Google Cloud Vision), speech recognition software (e.g., Google Speech-to-Text), large-scale language models (e.g., OpenAI® GPT-3®), sentiment analysis engines (e.g., Affectiva)
[0401] Initial Setup and Data Entry
[0402] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the basic information sent by the user, coordinates it with past conversation history data to perform initial setup, and generates optimal initial topics and suggestions (e.g., date plans or sales strategies). The user selects a suggested plan on the smartphone app, and the server arranges reservations and other preparations based on the selected plan.
[0403] Real-time data acquisition and analysis
[0404] The user wears an AR device (e.g., AR glasses or AR contact lenses), and the device captures video and audio in real time. The captured data is sent to a server, which analyzes it using image and voice recognition software. The analysis results are processed by an emotion engine, which analyzes the emotional state of the other person and the user themselves.
[0405] Generating and displaying advice
[0406] Based on the analysis results, the server uses a large-scale language model to generate optimal topics and behaviors. The generated advice is sent to the user's AR device and displayed as an overlay in the user's field of view. The user can refer to the displayed advice to proceed with the conversation.
[0407] Follow-up
[0408] After a date or business meeting ends, the user notifies the end of the event via a smartphone app. The server analyzes all data up to the end of the event and generates the next follow-up action (such as sending a message). The generated message is notified to the user, who can then send it after confirming it.
[0409] Specific examples
[0410] Specific examples of dating
[0411] 1. The day before the date
[0412] The user enters the other person's basic information into a smartphone app.
[0413] The server will suggest plans such as "lunch at a cafe" or "watching a movie."
[0414] The user selects "Watch a Movie" and the server reserves the tickets.
[0415] 2. The day of the date
[0416] The user puts on the AR glasses.
[0417] During the date, the AR glasses transmit video and audio to a server.
[0418] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[0419] A large-scale language model generates advice like, "Tell me about a movie you recently saw."
[0420] The advice is displayed in the user's field of view, and the user provides the topic.
[0421] 3. After the date
[0422] The server generates a message saying "Thanks for today, see you again soon!"
[0423] The message is notified to the smartphone app, and the user confirms it before sending it.
[0424] Specific examples of sales
[0425] 1. Before the sales visit
[0426] The user enters the other company's information and the name of the person in charge into the smartphone app.
[0427] The server references transaction history and related information to prepare initial proposals and topics.
[0428] 2. During a sales visit
[0429] The user puts on the AR contact lenses.
[0430] During business hours, the AR contact lenses capture video and audio and send it to a server.
[0431] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[0432] A large-scale language model generates the advice "Introduce and demonstrate the new product."
[0433] The advice is displayed in the user's field of view and the user demonstrates it.
[0434] 3. After business hours
[0435] The server generates a follow-up email to notify the user.
[0436] The user checks the email and then sends it.
[0437] Prompt Sentence Examples
[0438] "If the other person seems to be having fun, what topic should I bring up next?"
[0439] "What should I do when my partner seems a little cranky?"
[0440] "Suggest topics based on topics that your partner has shown interest in on previous dates."
[0441] In this way, the system based on this invention analyzes the emotions of the user and the other party in real time and suggests optimal topics and actions, thereby improving the quality of communication. The introduction of an emotion engine makes it easier for users to understand their own emotions, enabling more effective support.
[0442] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0443] Step 1:
[0444] The user launches the smartphone app. The user enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives this basic information and performs initial setup in conjunction with past conversation history data. Specifically, the server stores the entered name and hobbies of the other person in a database and compares them with past conversation history data. As an output, the most suitable initial topic and suggestions for the other person are generated based on the comparison results.
[0445] Step 2:
[0446] The user selects a proposed plan (e.g., a date plan or a sales strategy) on a smartphone app. The server receives the selected plan and arranges and prepares the necessary reservations based on that plan. At this time, the server connects to the reservation system and secures the necessary resources (e.g., movie tickets, restaurant reservations). As an output, the confirmed date plan or sales strategy is generated and notified to the user.
[0447] Step 3:
[0448] A user puts on an augmented reality (AR) device (e.g., AR glasses or AR contact lenses). The device captures video and audio in real time. This data is collected through the device's camera and microphone. The input real-time data is sent to a server, which prepares it for data analysis. Specifically, video data is sent to the server in streaming format, and audio data is transferred in real time. Data begins to be sent to the server as output.
[0449] Step 4:
[0450] The server analyzes the video data using image recognition software (e.g., Google Cloud Vision) and the audio data using voice recognition software (e.g., Google Speech-to-Text). In this analysis, the other person's facial expressions and tone of voice are extracted as important elements. The input is video and audio data transmitted in real time, and the output is data indicating the emotional state. Specifically, the server extracts facial expression features from the video data and extracts tone of voice and word features from the audio data.
[0451] Step 5:
[0452] The emotion engine analyzes the emotional state of the user and the other party in real time based on the analysis results. The input is the analysis results of image recognition and voice recognition, and the output is emotion analysis data. The emotion engine integrates this data and grasps the psychological state of the user and the other party in real time. Specifically, the emotion engine also monitors changes in the user's heart rate and breathing.
[0453] Step 6:
[0454] The server inputs the sentiment analysis data into a large-scale language model (e.g., OpenAI GPT-3) to generate optimal topics and behaviors. The input is the sentiment analysis data and provided context information, and the output is specific advice. By sending prompts to the large-scale language model, responses are generated to questions such as, "If the other person seems to be having fun, what topic should I bring up next?"
[0455] Step 7:
[0456] The server compiles the generated advice and sends it to the user's AR device. The input is advice generated by a large-scale language model, and the output is information overlaid on the user's field of view. Specifically, the server sends appropriate overlay information to the AR device, and the advice is displayed in the user's field of view.
[0457] Step 8:
[0458] The user proceeds with the dialogue while referring to the displayed advice. The input is the advice displayed on the AR device, and the output is the user's actions. The user interacts with the other party based on the advice, and the results are sent to the server as data acquired in real time. Specifically, the dialogue content and the other party's reactions are continuously acquired in real time.
[0459] Step 9:
[0460] After a date or business meeting ends, the user notifies the smartphone app that the event is over. The server analyzes all data up to the end and generates the follow-up action to be taken next. The input is all data up to the end, and the output is the follow-up action (e.g., sending a message). In concrete terms, the server comprehensively analyzes all the data obtained and determines the next step.
[0461] Step 10:
[0462] The server notifies the user of a follow-up message generated by the server. The user confirms and sends the message. The input is the message generated by the server, and the output is the message confirmed and sent by the user. Specifically, the user confirms the message in the smartphone app, edits or corrects it as necessary, and then taps the send button.
[0463] (Application example 2)
[0464] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0465] In conventional marketing activities, it has been difficult to instantly grasp the emotional reactions of clients and customers and propose optimal advertising strategies in real time. This has led to problems such as not being able to maximize the effectiveness of advertising and insufficient communication with clients. In particular, it has been difficult to provide effective feedback in situations where it is necessary to respond quickly to client reactions during a conversation.
[0466] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0467] In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through speech recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and a means for analyzing the emotional responses of conversation partners in real time during marketing activities and generating optimal advertising strategies and proposals. This makes it possible to analyze the emotions of clients or customers in real time and quickly generate optimal advertising strategies and proposals accordingly. This system is effective for marketers to receive instant feedback during conversations and promote effective communication with clients.
[0468] "Image recognition" is a technology that extracts and analyzes specific information and features from image data acquired using devices such as cameras.
[0469] "Means for estimating emotions" refers to technology that uses image recognition and voice recognition to analyze a person's facial expressions and tone of voice, and then estimates emotions based on the results.
[0470] "Speech recognition" is a technology that analyzes voice data acquired through devices such as microphones and extracts characteristics of words and voices.
[0471] A "large-scale language model" is a machine learning model trained using massive amounts of text data, and has the ability to understand and generate natural language.
[0472] "Means for generating optimal topics and behaviors" refers to a technology that uses a large-scale language model to generate suggestions for optimal topics and actions according to the situation.
[0473] A "user-worn augmented reality (AR) device" is a device worn by a user that can overlay digital information on the user's field of vision.
[0474] "Marketing activities" refers to various sales activities aimed at promoting products and services.
[0475] "Means for analyzing the emotional responses of a conversation partner in real time" refers to technology that instantly analyzes the emotions and responses of the other person during a conversation and provides that information.
[0476] An "advertising strategy" refers to a plan or method for effectively promoting a particular product or service.
[0477] The "means for generating suggestions" is a technology that suggests optimal actions and conversation content to users based on analyzed data.
[0478] The present invention provides a system for analyzing the emotions of a conversation partner in real time during marketing activities and for optimizing advertising strategies and proposals. The system is realized using image recognition, speech recognition, a large-scale language model, a wearable augmented reality (AR) device, and an emotion analysis engine. Specific embodiments for implementing the present invention are described below.
[0479] Overall system configuration
[0480] The system consists of the following main components:
[0481] 1. Client device (augmented reality (AR) device worn by the user)
[0482] 2. Server (responsible for major processing on the cloud)
[0483] 3. Emotion engine (analyzing the emotions of the person interacting with you and the user in real time)
[0484] 4. User interface (smartphone app or web interface)
[0485] Initial setup and preparation phase
[0486] First, the user launches the smartphone app and registers the other party's basic information (such as name, age, hobbies, and response to previous advertisements) in natural language. The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate an optimal initial advertising strategy and proposal for the other party and propose a marketing plan. The user selects the proposed plan and sends it to the server, completing preparations for marketing activities.
[0487] Real-time support during marketing activities
[0488] The user wears an augmented reality (AR) device, which captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the emotion of the person they are interacting with. At the same time, an emotion engine analyzes and monitors the user's emotional state in real time.
[0489] The server integrates the emotional information of the user and the conversation partner, analyzes the context of current marketing activities using a large-scale language model, and generates optimal advertising strategies and proposals. The generated proposals are displayed in the user's field of view, and the user continues the dialogue using them as reference. Furthermore, the emotion engine learns the user's past emotional data, analyzes their emotional tendencies, and makes proposals that correspond to predicted emotional changes.
[0490] Follow-up
[0491] After the marketing activity is completed, the user notifies the smartphone app. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending an email), and notifies the user. The notified message can be sent after the user confirms it.
[0492] Hardware and software used
[0493] Hardware:
[0494] Smartphone
[0495] camera
[0496] Augmented Reality (AR) Devices
[0497] software:
[0498] OpenCV: Image Recognition Library
[0499] dlib: face detection library
[0500] DeepFace: A library for emotion analysis
[0501] OpenAI API: Large-scale language models
[0502] Specific examples
[0503] During a meeting with an important client, a marketing manager at an advertising agency uses a smartphone camera to analyze the client's facial expressions. The manager and his colleagues then review the emotion analysis results to devise strategies and make optimal proposals in real time.
[0504] Prompt Sentence Examples
[0505] My client is expressing "delight" emotions. What advertising strategy should I suggest next?
[0506] My client is expressing "anger" emotions. How do I fix this next?
[0507] My client is expressing feelings of sadness. How should I follow up next?
[0508] As described above, the present invention provides a system that combines sentiment analysis and large-scale language models to provide optimal advertising strategies and proposals in real time for marketing activities.
[0509] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0510] Step 1:
[0511] The server receives basic information about the other party (such as name, age, hobbies, and response to previous advertisements) from the user's device (smartphone app). This allows the necessary information to be accumulated on the server. Natural language data entered on the user's device is sent to the server, where it is analyzed and saved.
[0512] Step 2:
[0513] The server references the received contact's basic information and past conversation history data to perform initial setup. It also generates and proposes an appropriate initial advertising strategy and marketing plan. The plan is generated using a large-scale language model, using the contact's basic information and past conversation history data as input.
[0514] Step 3:
[0515] The user selects a proposed plan and sends it to the server. The server determines a marketing activity plan based on the selected plan and notifies the user's device. The selected plan and related data are processed and stored on the server.
[0516] Step 4:
[0517] A user wears an augmented reality (AR) device and starts interacting with the client. The AR device captures video and audio in real time and sends them to the server. The video and audio data are sent to the server as input.
[0518] Step 5:
[0519] The server analyzes the received video data using OpenCV and dlib to perform facial recognition and emotion estimation (image recognition). At the same time, it analyzes the received audio data using DeepFace to estimate emotions (speech recognition). The results of the emotion analysis are generated and saved on the server.
[0520] Step 6:
[0521] The emotion analysis engine monitors the emotional state of the user and the person they are interacting with based on the analyzed emotion data. The monitoring data is created on the server.
[0522] Step 7:
[0523] It uses a large-scale language model to analyze the context of current marketing activities and generate optimal advertising strategies and proposals. Using prompt sentences, it generates specific advertising strategies and proposals and displays them in the user's field of view. The generated advertising strategies and proposals are displayed on the user's AR device.
[0524] Step 8:
[0525] The user visually confirms the suggestions through the AR device and applies them during the dialogue to communicate with the client. The user advances the conversation based on the generated strategies and suggestions.
[0526] Step 9:
[0527] After completing a marketing activity, the user notifies the server using the smartphone app.
[0528] Step 10:
[0529] The server comprehensively analyzes all log data and sentiment analysis data and generates the next follow-up action to be taken. The generated action and the content of the follow-up are notified to the user's device.
[0530] Step 11:
[0531] The user confirms the follow-up action (e.g., sending an email) and sends it as is if necessary. The follow-up action is executed from the user's device.
[0532] Through the above steps, the system of the present invention can analyze the emotions of the person in conversation in real time during marketing activities and make optimal advertising strategies and proposals.
[0533] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0534] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0535] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0536] [Second embodiment]
[0537] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0538] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0539] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0540] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0541] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0542] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0543] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0544] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0545] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0546] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0547] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0548] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0549] The present invention is a system for facilitating communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. The following describes how the system of the present invention is specifically implemented.
[0550] Overall system configuration
[0551] The system consists of the following main components:
[0552] 1. Client device (augmented reality (AR) device worn by the user)
[0553] 2. Server (responsible for major processing on the cloud)
[0554] 3. User interface (smartphone app or web interface)
[0555] Initial setup and preparation phase
[0556] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[0557] Real-time support during business hours
[0558] The user puts on AR glasses or AR contact lenses, and the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. The server integrates the emotion estimation results with the current situation and generates optimal advice using a large-scale language model. This allows the user to receive suggestions for appropriate topics and actions in real time. The advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[0559] Follow-up
[0560] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user.
[0561] Specific examples
[0562] In the case of a date
[0563] 1. The day before the date
[0564] The user enters the other person's basic information into a smartphone app.
[0565] The server suggests a plan of "lunch at Cafe A" and "watch movie B."
[0566] The user selects "Watch Movie B" and the server reserves the tickets.
[0567] 2. The day of the date
[0568] The user puts on the AR glasses.
[0569] During the date, the AR glasses transmit video and audio to a server.
[0570] The server infers that the other person is "interested" based on their facial expression and tone of voice.
[0571] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[0572] The advice is displayed in the user's field of view, and the user provides the topic.
[0573] 3. After the date
[0574] The server generates a message saying, "Thank you for today. It was fun!"
[0575] The message is notified to the smartphone app, and the user confirms it before sending it.
[0576] In the case of sales
[0577] 1. Before the sales visit
[0578] The user enters the other company's information and the name of the person in charge into the smartphone app.
[0579] The server prepares initial proposals and topics based on past transaction history and related information.
[0580] 2. During a sales visit
[0581] The user puts on the AR contact lenses.
[0582] During business hours, the AR contact lenses capture video and audio and send them to a server.
[0583] The server analyzes the facial expressions and tone of voice of the person in charge and determines that they are "interested."
[0584] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[0585] The advice is displayed in the user's field of view and the user demonstrates it.
[0586] 3. After the visit
[0587] The server generates a follow-up email.
[0588] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[0589] Thus, the present invention is a system that supports users in real time in choosing the best topics and actions to take during dates or business trips, thereby improving the quality of communication.
[0590] The processing flow will be explained below.
[0591] Initial setup and preparation phase
[0592] Step 1:
[0593] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[0594] Step 2:
[0595] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received data, and retrieves matching data.
[0596] Step 3:
[0597] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[0598] Step 4:
[0599] The user checks the proposed date plans and sales strategies, selects one, and sends the selection information to the server.
[0600] Step 5:
[0601] The server makes the necessary reservations (restaurants, cinemas, etc.) based on the user's selection, and notifies the user's app that the reservation has been completed.
[0602] Real-time support during business hours
[0603] Step 1:
[0604] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[0605] Step 2:
[0606] The device uses its sensors to capture video and audio in real time, and sends the captured data to a server.
[0607] Step 3:
[0608] The server analyzes the received video data using an image recognition algorithm to infer emotions from facial expressions, and analyzes the audio data using a voice recognition algorithm to infer emotions from the tone and patterns of the voice.
[0609] Step 4:
[0610] The server sends the emotion estimation results to the user's smartphone app, which simultaneously analyzes the emotion information and the current dating or sales context using a large-scale language model to generate optimal advice.
[0611] Step 5:
[0612] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[0613] Follow-up
[0614] Step 1:
[0615] The user enters into the app that a date or business has ended, and the device sends all data up to the end to the server.
[0616] Step 2:
[0617] The server analyzes the data at the end, generates the next follow-up action to be taken, and notifies the smartphone app with the generated recommendation message.
[0618] Step 3:
[0619] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[0620] Example 1
[0621] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0622] In modern society, it is difficult to grasp appropriate topics and actions in real time in communication situations such as dating or business. In particular, accurately understanding one's own emotions and the emotions of others and acting accordingly is extremely important, but not easy to achieve. Conventional systems lack the ability to assess the other person's emotions and provide appropriate advice, resulting in a decline in the quality of communication. New technology is needed to solve these problems.
[0623] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0624] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for acquiring video and audio in real time and transmitting them to the server, and means for analyzing all data up to the end and generating the next action. This allows the user to receive suggestions for appropriate topics and actions in real time, enabling high-quality communication even during dates or business trips.
[0625] "Image recognition" is a technology that analyzes image data acquired using cameras and sensors to identify objects, people, features, etc.
[0626] "Emotion estimation" is a technology that determines the other person's psychological state and emotions by analyzing image and audio data.
[0627] "Speech recognition" is a technology that analyzes voice data acquired using devices such as microphones and identifies what a person is saying and the characteristics of their voice.
[0628] A "large-scale language model" is an artificial intelligence model that learns from large amounts of text data to generate and understand natural language.
[0629] An "augmented reality (AR) device" is a device that displays digital information superimposed on real-world visual information.
[0630] "Real-time" means processing acquired data immediately without delay and providing results instantly.
[0631] "Capturing video and audio" means collecting visual and audio information from the surroundings using devices such as cameras and microphones.
[0632] A "server" is a central computer system that provides services over a network.
[0633] "Analyzing all data" means extracting and analyzing information for a specific purpose from all acquired data.
[0634] "Generating the next action" means proposing the next action or response that the user should take based on the analysis results.
[0635] This invention is a system for facilitating smooth communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. Below, we will explain how to specifically implement the system of this invention.
[0636] Hardware and software used
[0637] Client device: An augmented reality (AR) device worn by a user (e.g., AR glasses, AR contact lenses).
[0638] Server: A computer system that handles major processing on the cloud.
[0639] User interface: Smartphone app and web interface.
[0640] Initial setup and preparation phase
[0641] 1. Enter information
[0642] The user starts the smartphone app and enters the other person's basic information (name, age, hobbies, previous conversations), which is then received by the server and stored in a database.
[0643] 2. Information Reception and Analysis
[0644] The server analyzes the received information and generates initial topics and suggestions based on past conversation history data. The server then uses a generative AI model to propose date plans and sales strategies.
[0645] 3. Plan proposal
[0646] The server proposes specific plans to the user, such as "lunch at Cafe A" or "watch movie B." The user selects one of the proposed plans and sends it to the server, which then arranges the reservations required for the selected plan.
[0647] Example prompt sentence:
[0648] Based on the other person's basic information, propose a date plan (e.g., lunch at Cafe A, watching Movie B).
[0649] Real-time support during business hours
[0650] 1. Wearing the device
[0651] Users wear AR glasses or AR contact lenses, and these devices capture video and audio in real time.
[0652] 2. Data Acquisition and Transmission
[0653] The device sends the captured video and audio to a server, which analyzes them and uses image and voice recognition technology to estimate the other person's emotions.
[0654] 3. Emotion Estimation and Advice Generation
[0655] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[0656] 4. Displaying Advice
[0657] The server sends the generated advice to the user's device, which then displays it in the user's field of view. For example, if the person speaking is smiling, the advice displayed might be, "Try talking about a movie you recently saw."
[0658] Example prompt sentence:
[0659] Based on the video and audio captured in real time, estimate the other person's emotions (e.g., interest) and generate appropriate topics of conversation.
[0660] Follow-up
[0661] 1. Notice of Termination
[0662] The user notifies the smartphone app that a date or business meeting has ended. The server collects all data up to that point and begins analyzing it.
[0663] 2. Suggested next actions
[0664] The server generates the next action to be taken. For example, the server generates a follow-up message such as "Thank you for today. I had a great time!" and notifies the user's smartphone app. The follow-up message is sent when the user confirms and sends it.
[0665] Example prompt sentence:
[0666] Generate a follow-up message to send after the date ends (e.g., "Thank you for today. I had a great time!").
[0667] In this way, the system of the present invention can support users in real time to choose the most appropriate topic and action in any situation, thereby realizing high-quality communication even during dates or business meetings.
[0668] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0669] Step 1:
[0670] Enter information
[0671] The user launches the smartphone app and enters the other person's basic information (name, age, hobbies, and previous conversation details).
[0672] Input: The user enters the following information into the input form: "Name = Tanaka Taro, Age = 28, Hobby = Reading, Previous conversation content = We talked about movies."
[0673] Output: The data entered by the user is saved in the app and sent to the server.
[0674] Step 2:
[0675] Information reception and analysis
[0676] The server receives the basic information sent by the user.
[0677] The server analyzes the received data and stores it in a database. By linking it with past conversation history data, it generates optimal initial topics and suggestions.
[0678] Input: Basic information data submitted by the user.
[0679] Data processing: Obtain past conversation data from the database and analyze the other person's interests.
[0680] Output: Generate a list of initial topics and suggestions.
[0681] Step 3:
[0682] Plan proposal
[0683] The server uses the generated AI model to propose specific date plans and sales strategies.
[0684] Example: Generate a plan that includes "Lunch at Cafe X" and "Watch Movie Y."
[0685] Input: A candidate list and a prompt sentence for the generative AI model.
[0686] Data processing: The AI model generates a date plan based on the prompt.
[0687] Output: A list of specific plans proposed to the user.
[0688] Step 4:
[0689] Plan selection and reservation arrangements
[0690] The user selects a plan suggested by the smartphone app. For example, they select "Watch Movie Y."
[0691] The server then arranges the necessary reservations for the selected plan, in this case booking movie tickets online.
[0692] Input: The user's selected plan.
[0693] Data processing: Connect to the online reservation system based on the selected plan and complete the reservation.
[0694] Output: Reservation confirmation information is sent from the server and notified to the user.
[0695] Step 5:
[0696] Wearing the device
[0697] The user puts on AR glasses or AR contact lenses before starting a date or business trip.
[0698] Input: The act of a user putting on a device.
[0699] Output: The AR device is initialized and begins capturing video and audio in real time.
[0700] Step 6:
[0701] Acquiring and Sending Data
[0702] The device captures video and audio in real time and transmits them to the server.
[0703] Input: Surrounding video and audio information.
[0704] Data processing: The device acquires the information, encodes it, and uploads it to the server.
[0705] Output: The data stream that the server receives in real time.
[0706] Step 7:
[0707] Emotion estimation and advice generation
[0708] The server analyzes the received video and audio and uses image and voice recognition technology to estimate the other person's emotions. For example, it detects the other person's smile and infers that they are "interested."
[0709] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[0710] Input: Received video and audio data.
[0711] Data processing: Analyzes video and audio to recognize the other person's emotions. A large-scale language model generates dialogue advice based on the recognition results.
[0712] Output: The generated advice is displayed in the user's field of view.
[0713] Step 8:
[0714] Displaying Advice
[0715] The server transmits the generated advice to the user's terminal.
[0716] The device displays advice in the user's field of view, such as "Tell me about a movie you recently saw."
[0717] Input: The generated advice.
[0718] Data processing: Converting advice into a visually displayable format.
[0719] Output: An advisory message that is displayed in the user's field of view.
[0720] Step 9:
[0721] Notice of Termination
[0722] The smartphone app notifies the user when a date or business meeting has ended.
[0723] Input: User clicks the "Exit" button.
[0724] Output: A completion notice is sent to the server.
[0725] Step 10:
[0726] Data analysis
[0727] The server collects all data up to the end point and begins analyzing it, for example, analyzing the content of the conversation that day and the changes in the other person's emotions.
[0728] Input: Termination notice and all collected data.
[0729] Data processing: Analyze conversation content and emotional data to extract trends.
[0730] Output: A report of the analysis results.
[0731] Step 11:
[0732] Suggested next actions
[0733] The server generates the next action to be taken, for example, a follow-up message such as "Thank you for today. I had a great time!"
[0734] Input: Analysis results.
[0735] Data processing: The next action is generated using a generative AI model.
[0736] Output: The actions and messages generated.
[0737] Step 12:
[0738] Notification and confirmation
[0739] The server generates a message and sends it to the user's smartphone app. The user confirms it and sends the message.
[0740] Input: The generated message.
[0741] Data processing: Notify the message and ask for confirmation.
[0742] Output: The message that is confirmed by the user and sent to the other party.
[0743] (Application example 1)
[0744] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0745] Existing customer service support systems make it difficult for staff to provide optimal product recommendations and customer service methods for each individual customer in a physical store. This makes it difficult to provide effective customer service to increase customer satisfaction, and it is difficult to increase store sales. In addition, because the system relies on the customer service skills of the staff, it is difficult for new or inexperienced staff to provide effective customer service.
[0746] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0747] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for analyzing the customer's facial expressions and acquiring facial emotion data, means for analyzing voice data and converting it into text, means for generating appropriate product proposals by combining the text data and the facial emotion data, and means for synthesizing the generated product proposals and notifying the staff. This enables the staff to grasp the customer's emotions and conversation content in real time and immediately receive optimal product proposals and customer service methods.
[0748] "Image recognition" is a technology that detects and analyzes the characteristics of specific objects or people from image data.
[0749] "Emotion estimation" means estimating a person's emotional state (e.g., joy, anger, sadness, etc.) based on image and audio data.
[0750] "Speech recognition" is a technology that analyzes voice data and converts it into text data.
[0751] A "large-scale language model" is an artificial intelligence model trained on vast amounts of text data that is capable of understanding and generating natural language.
[0752] An "augmented reality (AR) device" is a device that displays computer graphics overlaid on images of the real world.
[0753] "Analyzing facial expressions to obtain facial emotion data" means using image recognition technology to extract emotional information from a person's facial expressions and obtain that data.
[0754] "Analyzing voice data and converting it into text" means converting voice data into text data using voice recognition technology.
[0755] "Generating appropriate product proposals by combining text data and facial emotion data" means automatically generating optimal product proposals for customers using text data and facial emotion data.
[0756] "Synthesize voice and notify staff" means synthesizing the generated product suggestions and advice into voice and conveying it to staff.
[0757] This invention provides a system that makes full use of image recognition, voice recognition, and large-scale language models to support customer service in brick-and-mortar stores. Specific embodiments for implementing this system are described below.
[0758] System configuration
[0759] The system consists of the following main components:
[0760] 1. Client devices (including smart glasses and head-mounted displays)
[0761] 2. Server (located on the cloud and responsible for main processing)
[0762] 3. User interface (smartphone app or web interface)
[0763] Program Overview
[0764] Emotion estimation using image recognition
[0765] The client device captures the customer's facial expressions in real time while serving customers in a physical store and sends the image data to a server. The server then uses image recognition technology such as Amazon Rekognition to analyze the customer's emotions (e.g., joy, anger, sadness, etc.) and obtains facial emotion data.
[0766] Emotion estimation using speech recognition
[0767] The client device records the conversation with the customer and sends the audio data to the server. The server then converts the audio data into text data using voice recognition technology such as the Google Speech-to-Text API. This text data is then used to estimate emotions based on the conversation content.
[0768] Generating optimal suggestions using large-scale language models
[0769] The server uses a large-scale language model (e.g., GPT-4) to generate optimal product suggestions and customer service methods based on text data and facial emotion data. The generated suggestions are input as prompt sentences as follows:
[0770] Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation: "What products are you interested in recently?"
[0771] Display suggestions and speech synthesis
[0772] The generated suggestions are notified to staff using voice synthesis technology (such as pyttsx3). The user (staff member) receives this information in real time through smart glasses or a head-mounted display attached to the client device and can take appropriate action.
[0773] Specific examples
[0774] Example 1: Proposing a new product
[0775] When a customer asks a staff member, "What products do you recommend?", the system converts the voice data into text data and analyzes facial expressions to infer that the customer is "interested." Based on this information, a large-scale language model generates a suggestion such as "Recommend a newly arrived product," and notifies the staff member via voice synthesis.
[0776] Example 2: Follow-up suggestions
[0777] After the staff member has finished serving the customer, the system analyzes the acquired data and generates an appropriate follow-up action (e.g., sending a thank-you message) and notifies the user interface. The user can then confirm the suggestion and take action if necessary.
[0778] As a result, this system allows staff to understand customers' emotions and conversation content in real time, enabling them to instantly receive optimal product suggestions and customer service methods.
[0779] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0780] Step 1:
[0781] The user wears the client terminal.
[0782] How it works: The customer wears smart glasses or a head-mounted display, which allows the device's camera and microphone to capture the customer's video and audio in real time.
[0783] Input: The client device worn by the user.
[0784] Output: Real-time video and audio capture.
[0785] Step 2:
[0786] The terminal captures the customer's video and audio and sends them to the server.
[0787] How it works: The device's camera captures the customer's facial expressions, and the microphone records the conversation. This data is then sent to a cloud server.
[0788] Input: Video and audio data acquired by the device.
[0789] Output: Video and audio data sent to the server.
[0790] Step 3:
[0791] The server uses image recognition technology to analyze the customer's facial expressions and obtain facial emotion data.
[0792] How it works: The server runs an image recognition algorithm such as Amazon Rekognition to estimate the customer's emotions from their facial expressions. For example, the results can be "happy" or "interested" based on the facial expression.
[0793] Input: Video data sent to the server.
[0794] Output: Facial emotion data (e.g., "HAPPY", confidence level 95.0%).
[0795] Step 4:
[0796] The server uses voice recognition technology to convert the voice data into text data.
[0797] What happens: The server runs a speech recognition algorithm, such as the Google Speech-to-Text API, to convert the conversation into text. For example, it might generate something like, "What products are you interested in these days?"
[0798] Input: The audio data sent to the server.
[0799] Output: Text data (e.g., "What products have you been interested in recently?").
[0800] Step 5:
[0801] The server combines text data and facial emotion data to input prompts into a large-scale language model to generate optimal product suggestions.
[0802] Specific operation: The server uses a large-scale language model (e.g., GPT-4) to generate a prompt sentence and then generates appropriate product suggestions based on it. An example of a prompt sentence is "Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation content: 'What products are you interested in recently?'"
[0803] Input: Text data and facial emotion data.
[0804] Output: Product suggestions (e.g. "Recommend new products").
[0805] Step 6:
[0806] The product suggestions generated by the server are notified to staff using voice synthesis technology.
[0807] Specific operation: The server uses a speech synthesis algorithm such as pyttsx3 to convert the generated product suggestions into voice and send it to the client terminal. The user (staff member) receives the suggestions by voice and can respond appropriately.
[0808] Input: Server-generated product suggestions.
[0809] Output: A voice-synthesized product suggestion notification.
[0810] Step 7:
[0811] Based on the user-generated proposals, appropriate products are proposed to the customer.
[0812] Specific operation: The user (staff member) uses the product proposal notified by voice from the client terminal as a reference to explain and propose new products to the customer.
[0813] Input: Speech-synthesized product suggestions.
[0814] Output: Product recommendations to customers.
[0815] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0816] This invention is a system that analyzes the emotions of users and their partners in real time during dates and business meetings, and suggests optimal topics and behaviors. This system achieves highly accurate emotion analysis by combining image recognition, speech recognition technology, large-scale language models, and an emotion engine.
[0817] Overall system configuration
[0818] The system consists of the following main components:
[0819] 1. Client device (augmented reality (AR) device worn by the user)
[0820] 2. Server (responsible for major processing on the cloud)
[0821] 3. Emotion engine (analyzing user and other party emotions in real time)
[0822] 4. User interface (smartphone app or web interface)
[0823] Initial setup and preparation phase
[0824] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[0825] Real-time support during business hours
[0826] When a user wears AR glasses or AR contact lenses, the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. At the same time, an emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[0827] The server integrates the emotional information of both the user and the other party, analyzes the current dating or sales context using a large-scale language model, and generates optimal advice. The generated advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[0828] Furthermore, the emotion engine learns from the user's past emotional data, analyzes their emotional tendencies, and provides advice in response to predicted emotional changes. This makes it easier for the user to understand their own emotional state and continue the conversation calmly.
[0829] Follow-up
[0830] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user. The notified message can be sent after the user confirms it.
[0831] Specific examples
[0832] In the case of a date
[0833] 1. The day before the date
[0834] The user enters the other person's basic information into a smartphone app.
[0835] The server suggests plans such as "lunch at Cafe A" or "watching Movie B."
[0836] The user selects "Watch Movie B" and the server reserves the tickets.
[0837] 2. The day of the date
[0838] The user puts on the AR glasses.
[0839] During the date, the AR glasses transmit video and audio to a server.
[0840] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[0841] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[0842] The advice is displayed in the user's field of view, and the user provides the topic.
[0843] 3. After the date
[0844] The server generates a message saying, "Thank you for today. It was fun!"
[0845] The message is notified to the smartphone app, and the user confirms it before sending it.
[0846] In the case of sales
[0847] 1. Before the sales visit
[0848] The user enters the other company's information and the name of the person in charge into the smartphone app.
[0849] The server prepares initial proposals and topics based on past transaction history and related information.
[0850] 2. During a sales visit
[0851] The user puts on the AR contact lenses.
[0852] During business hours, the AR contact lenses capture video and audio and send them to a server.
[0853] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[0854] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[0855] The advice is displayed in the user's field of view and the user demonstrates it.
[0856] 3. After the visit
[0857] The server generates a follow-up email.
[0858] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[0859] In this way, the present invention is a system that improves the quality of communication by analyzing the emotions of the user and the other party in real time and suggesting optimal topics and actions. The introduction of an emotion engine makes it easier to understand the user's own emotions, enabling more effective support.
[0860] The processing flow will be explained below.
[0861] Initial setup and preparation phase
[0862] Step 1:
[0863] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[0864] Step 2:
[0865] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received information, and retrieves matching data.
[0866] Step 3:
[0867] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[0868] Step 4:
[0869] The user checks the proposed date plans and sales strategies and selects one. The selected information is sent to the server.
[0870] Step 5:
[0871] The server makes the necessary reservations (for restaurants, cinemas, etc.) based on the user's selection, and notifies the user's smartphone app that the reservation has been completed.
[0872] Real-time support during business hours
[0873] Step 1:
[0874] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[0875] Step 2:
[0876] The device uses its sensors to capture video and audio in real time, and the captured data is sent to a server.
[0877] Step 3:
[0878] The server analyzes the video data it receives using an image recognition algorithm to infer emotions from facial expressions, and it also analyzes audio data using a voice recognition algorithm to infer emotions from the tone and patterns of voice.
[0879] Step 4:
[0880] The server creates the other person's emotional information based on the emotion estimation results. At the same time, the emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[0881] Step 5:
[0882] The server integrates the emotional information of both the user and the other party and generates optimal advice using a large-scale language model. The advice is then displayed in the user's field of view.
[0883] Step 6:
[0884] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[0885] Follow-up
[0886] Step 1:
[0887] The user enters into the app that a date or business trip has ended, and the device sends all data up to the end to the server.
[0888] Step 2:
[0889] The server analyzes the data at the end and generates a recommended action to be taken next, which is then sent to the user's smartphone app.
[0890] Step 3:
[0891] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[0892] Example 2
[0893] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0894] In conventional dating and sales situations, it has been difficult for users to accurately read the other person's emotions and decide on appropriate topics and actions. Furthermore, it has been difficult for users to manage their own emotions while having an effective conversation. This has led to a decline in the quality of communication and a decrease in the success rate of dates and sales.
[0895] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through voice recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and an emotion engine for analyzing the user's emotions in real time. This configuration enables the user to grasp the emotions of others and themselves in real time and take optimal actions and responses.
[0896] "Image recognition" is the technology that extracts information from digital images and videos and identifies specific patterns and features.
[0897] "Speech recognition" is a technology that analyzes voice data and converts spoken words into text data.
[0898] A "large-scale language model" is an artificial intelligence model trained on large amounts of text data, and is a model that excels at understanding and generating natural language.
[0899] An "augmented reality (AR) device" is a device that overlays digital information onto a user's real-world field of vision. Examples include AR glasses and AR contact lenses.
[0900] The "emotion engine" is a system that analyzes the emotional state of the user and the other person in real time and provides appropriate feedback and advice based on the results.
[0901] A "date plan" refers to a proposal or action plan for systematically preparing the progress of a date.
[0902] A "sales strategy" is a plan for effective contact methods and presentations in sales activities.
[0903] "Real-time data" is data that represents ongoing events or conditions and should be processed immediately.
[0904] "Follow-up actions" are subsequent actions that should be taken after a date or business meeting, with the purpose of maintaining and strengthening the relationship with the other person.
[0905] "Initial settings" refers to the setup work required before starting to use the system, in which basic information about the user and the other party is entered and preparations required for system operation are made.
[0906] The system based on this invention analyzes the emotions of the user and the other party in real time during dating or business, and suggests optimal topics and behaviors. This system includes the following main components: image recognition, speech recognition, a large-scale language model, an augmented reality (AR) device, and an emotion engine.
[0907] Hardware and software used
[0908] Hardware: Augmented reality (AR) glasses, augmented reality (AR) contact lenses, smartphones
[0909] Software: Image recognition software (e.g., Google Cloud Vision), speech recognition software (e.g., Google Speech-to-Text), large-scale language models (e.g., OpenAI GPT-3), sentiment analysis engines (e.g., Affectiva)
[0910] Initial Setup and Data Entry
[0911] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the basic information sent by the user, coordinates it with past conversation history data to perform initial setup, and generates optimal initial topics and suggestions (e.g., date plans or sales strategies). The user selects a suggested plan on the smartphone app, and the server arranges reservations and other preparations based on the selected plan.
[0912] Real-time data acquisition and analysis
[0913] The user wears an AR device (e.g., AR glasses or AR contact lenses), and the device captures video and audio in real time. The captured data is sent to a server, which analyzes it using image and voice recognition software. The analysis results are processed by an emotion engine, which analyzes the emotional state of the other person and the user themselves.
[0914] Generating and displaying advice
[0915] Based on the analysis results, the server uses a large-scale language model to generate optimal topics and behaviors. The generated advice is sent to the user's AR device and displayed as an overlay in the user's field of view. The user can refer to the displayed advice to proceed with the conversation.
[0916] Follow-up
[0917] After a date or business meeting ends, the user notifies the end of the event via a smartphone app. The server analyzes all data up to the end of the event and generates the next follow-up action (such as sending a message). The generated message is notified to the user, who can then send it after confirming it.
[0918] Specific examples
[0919] Specific examples of dating
[0920] 1. The day before the date
[0921] The user enters the other person's basic information into a smartphone app.
[0922] The server will suggest plans such as "lunch at a cafe" or "watching a movie."
[0923] The user selects "Watch a Movie" and the server reserves the tickets.
[0924] 2. The day of the date
[0925] The user puts on the AR glasses.
[0926] During the date, the AR glasses transmit video and audio to a server.
[0927] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[0928] A large-scale language model generates advice like, "Tell me about a movie you recently saw."
[0929] The advice is displayed in the user's field of view, and the user provides the topic.
[0930] 3. After the date
[0931] The server generates a message saying "Thanks for today, see you again soon!"
[0932] The message is notified to the smartphone app, and the user confirms it before sending it.
[0933] Specific examples of sales
[0934] 1. Before the sales visit
[0935] The user enters the other company's information and the name of the person in charge into the smartphone app.
[0936] The server references transaction history and related information to prepare initial proposals and topics.
[0937] 2. During a sales visit
[0938] The user puts on the AR contact lenses.
[0939] During business hours, the AR contact lenses capture video and audio and send it to a server.
[0940] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[0941] A large-scale language model generates the advice "Introduce and demonstrate the new product."
[0942] The advice is displayed in the user's field of view and the user demonstrates it.
[0943] 3. After business hours
[0944] The server generates a follow-up email to notify the user.
[0945] The user checks the email and then sends it.
[0946] Prompt Sentence Examples
[0947] "If the other person seems to be having fun, what topic should I bring up next?"
[0948] "What should I do when my partner seems a little cranky?"
[0949] "Suggest topics based on topics that your partner has shown interest in on previous dates."
[0950] In this way, the system based on this invention analyzes the emotions of the user and the other party in real time and suggests optimal topics and actions, thereby improving the quality of communication. The introduction of an emotion engine makes it easier for users to understand their own emotions, enabling more effective support.
[0951] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0952] Step 1:
[0953] The user launches the smartphone app. The user enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives this basic information and performs initial setup in conjunction with past conversation history data. Specifically, the server stores the entered name and hobbies of the other person in a database and compares them with past conversation history data. As an output, the most suitable initial topic and suggestions for the other person are generated based on the comparison results.
[0954] Step 2:
[0955] The user selects a proposed plan (e.g., a date plan or a sales strategy) on a smartphone app. The server receives the selected plan and arranges and prepares the necessary reservations based on that plan. At this time, the server connects to the reservation system and secures the necessary resources (e.g., movie tickets, restaurant reservations). As an output, the confirmed date plan or sales strategy is generated and notified to the user.
[0956] Step 3:
[0957] A user puts on an augmented reality (AR) device (e.g., AR glasses or AR contact lenses). The device captures video and audio in real time. This data is collected through the device's camera and microphone. The input real-time data is sent to a server, which prepares it for data analysis. Specifically, video data is sent to the server in streaming format, and audio data is transferred in real time. Data begins to be sent to the server as output.
[0958] Step 4:
[0959] The server analyzes the video data using image recognition software (e.g., Google Cloud Vision) and the audio data using voice recognition software (e.g., Google Speech-to-Text). In this analysis, the other person's facial expressions and tone of voice are extracted as important elements. The input is video and audio data transmitted in real time, and the output is data indicating the emotional state. Specifically, the server extracts facial expression features from the video data and extracts tone of voice and word features from the audio data.
[0960] Step 5:
[0961] The emotion engine analyzes the emotional state of the user and the other party in real time based on the analysis results. The input is the analysis results of image recognition and voice recognition, and the output is emotion analysis data. The emotion engine integrates this data and grasps the psychological state of the user and the other party in real time. Specifically, the emotion engine also monitors changes in the user's heart rate and breathing.
[0962] Step 6:
[0963] The server inputs the sentiment analysis data into a large-scale language model (e.g., OpenAI GPT-3) to generate optimal topics and behaviors. The input is the sentiment analysis data and provided context information, and the output is specific advice. By sending prompts to the large-scale language model, responses are generated to questions such as, "If the other person seems to be having fun, what topic should I bring up next?"
[0964] Step 7:
[0965] The server compiles the generated advice and sends it to the user's AR device. The input is advice generated by a large-scale language model, and the output is information overlaid on the user's field of view. Specifically, the server sends appropriate overlay information to the AR device, and the advice is displayed in the user's field of view.
[0966] Step 8:
[0967] The user proceeds with the dialogue while referring to the displayed advice. The input is the advice displayed on the AR device, and the output is the user's actions. The user interacts with the other party based on the advice, and the results are sent to the server as data acquired in real time. Specifically, the dialogue content and the other party's reactions are continuously acquired in real time.
[0968] Step 9:
[0969] After a date or business meeting ends, the user notifies the smartphone app that the event is over. The server analyzes all data up to the end and generates the follow-up action to be taken next. The input is all data up to the end, and the output is the follow-up action (e.g., sending a message). In concrete terms, the server comprehensively analyzes all the data obtained and determines the next step.
[0970] Step 10:
[0971] The server notifies the user of a follow-up message generated by the server. The user confirms and sends the message. The input is the message generated by the server, and the output is the message confirmed and sent by the user. Specifically, the user confirms the message in the smartphone app, edits or corrects it as necessary, and then taps the send button.
[0972] (Application example 2)
[0973] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0974] In conventional marketing activities, it has been difficult to instantly grasp the emotional reactions of clients and customers and propose optimal advertising strategies in real time. This has led to problems such as not being able to maximize the effectiveness of advertising and insufficient communication with clients. In particular, it has been difficult to provide effective feedback in situations where it is necessary to respond quickly to client reactions during a conversation.
[0975] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0976] In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through speech recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and a means for analyzing the emotional responses of conversation partners in real time during marketing activities and generating optimal advertising strategies and proposals. This makes it possible to analyze the emotions of clients or customers in real time and quickly generate optimal advertising strategies and proposals accordingly. This system is effective for marketers to receive instant feedback during conversations and promote effective communication with clients.
[0977] "Image recognition" is a technology that extracts and analyzes specific information and features from image data acquired using devices such as cameras.
[0978] "Means for estimating emotions" refers to technology that uses image recognition and voice recognition to analyze a person's facial expressions and tone of voice, and then estimates emotions based on the results.
[0979] "Speech recognition" is a technology that analyzes voice data acquired through devices such as microphones and extracts characteristics of words and voices.
[0980] A "large-scale language model" is a machine learning model trained using massive amounts of text data, and has the ability to understand and generate natural language.
[0981] "Means for generating optimal topics and behaviors" refers to a technology that uses a large-scale language model to generate suggestions for optimal topics and actions according to the situation.
[0982] A "user-worn augmented reality (AR) device" is a device worn by a user that can overlay digital information on the user's field of vision.
[0983] "Marketing activities" refers to various sales activities aimed at promoting products and services.
[0984] "Means for analyzing the emotional responses of a conversation partner in real time" refers to technology that instantly analyzes the emotions and responses of the other person during a conversation and provides that information.
[0985] An "advertising strategy" refers to a plan or method for effectively promoting a particular product or service.
[0986] The "means for generating suggestions" is a technology that suggests optimal actions and conversation content to users based on analyzed data.
[0987] The present invention provides a system for analyzing the emotions of a conversation partner in real time during marketing activities and for optimizing advertising strategies and proposals. The system is realized using image recognition, speech recognition, a large-scale language model, a wearable augmented reality (AR) device, and an emotion analysis engine. Specific embodiments for implementing the present invention are described below.
[0988] Overall system configuration
[0989] The system consists of the following main components:
[0990] 1. Client device (augmented reality (AR) device worn by the user)
[0991] 2. Server (responsible for major processing on the cloud)
[0992] 3. Emotion engine (analyzing the emotions of the person interacting with you and the user in real time)
[0993] 4. User interface (smartphone app or web interface)
[0994] Initial setup and preparation phase
[0995] First, the user launches the smartphone app and registers the other party's basic information (such as name, age, hobbies, and response to previous advertisements) in natural language. The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate an optimal initial advertising strategy and proposal for the other party and propose a marketing plan. The user selects the proposed plan and sends it to the server, completing preparations for marketing activities.
[0996] Real-time support during marketing activities
[0997] The user wears an augmented reality (AR) device, which captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the emotion of the person they are interacting with. At the same time, an emotion engine analyzes and monitors the user's emotional state in real time.
[0998] The server integrates the emotional information of the user and the conversation partner, analyzes the context of current marketing activities using a large-scale language model, and generates optimal advertising strategies and proposals. The generated proposals are displayed in the user's field of view, and the user continues the dialogue using them as reference. Furthermore, the emotion engine learns the user's past emotional data, analyzes their emotional tendencies, and makes proposals that correspond to predicted emotional changes.
[0999] Follow-up
[1000] After the marketing activity is completed, the user notifies the smartphone app. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending an email), and notifies the user. The notified message can be sent after the user confirms it.
[1001] Hardware and software used
[1002] Hardware:
[1003] Smartphone
[1004] camera
[1005] Augmented Reality (AR) Devices
[1006] software:
[1007] OpenCV: Image Recognition Library
[1008] dlib: face detection library
[1009] DeepFace: A library for emotion analysis
[1010] OpenAI API: Large-scale language models
[1011] Specific examples
[1012] During a meeting with an important client, a marketing manager at an advertising agency uses a smartphone camera to analyze the client's facial expressions. The manager and his colleagues then review the emotion analysis results to devise strategies and make optimal proposals in real time.
[1013] Prompt Sentence Examples
[1014] My client is expressing "delight" emotions. What advertising strategy should I suggest next?
[1015] My client is expressing "anger" emotions. How do I fix this next?
[1016] My client is expressing feelings of sadness. How should I follow up next?
[1017] As described above, the present invention provides a system that combines sentiment analysis and large-scale language models to provide optimal advertising strategies and proposals in real time for marketing activities.
[1018] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1019] Step 1:
[1020] The server receives basic information about the other party (such as name, age, hobbies, and response to previous advertisements) from the user's device (smartphone app). This allows the necessary information to be accumulated on the server. Natural language data entered on the user's device is sent to the server, where it is analyzed and saved.
[1021] Step 2:
[1022] The server references the received contact's basic information and past conversation history data to perform initial setup. It also generates and proposes an appropriate initial advertising strategy and marketing plan. The plan is generated using a large-scale language model, using the contact's basic information and past conversation history data as input.
[1023] Step 3:
[1024] The user selects a proposed plan and sends it to the server. The server determines a marketing activity plan based on the selected plan and notifies the user's device. The selected plan and related data are processed and stored on the server.
[1025] Step 4:
[1026] A user wears an augmented reality (AR) device and starts interacting with the client. The AR device captures video and audio in real time and sends them to the server. The video and audio data are sent to the server as input.
[1027] Step 5:
[1028] The server analyzes the received video data using OpenCV and dlib to perform facial recognition and emotion estimation (image recognition). At the same time, it analyzes the received audio data using DeepFace to estimate emotions (speech recognition). The results of the emotion analysis are generated and saved on the server.
[1029] Step 6:
[1030] The emotion analysis engine monitors the emotional state of the user and the person they are interacting with based on the analyzed emotion data. The monitoring data is created on the server.
[1031] Step 7:
[1032] It uses a large-scale language model to analyze the context of current marketing activities and generate optimal advertising strategies and proposals. Using prompt sentences, it generates specific advertising strategies and proposals and displays them in the user's field of view. The generated advertising strategies and proposals are displayed on the user's AR device.
[1033] Step 8:
[1034] The user visually confirms the suggestions through the AR device and applies them during the dialogue to communicate with the client. The user advances the conversation based on the generated strategies and suggestions.
[1035] Step 9:
[1036] After completing a marketing activity, the user notifies the server using the smartphone app.
[1037] Step 10:
[1038] The server comprehensively analyzes all log data and sentiment analysis data and generates the next follow-up action to be taken. The generated action and the content of the follow-up are notified to the user's device.
[1039] Step 11:
[1040] The user confirms the follow-up action (e.g., sending an email) and sends it as is if necessary. The follow-up action is executed from the user's device.
[1041] Through the above steps, the system of the present invention can analyze the emotions of the person in conversation in real time during marketing activities and make optimal advertising strategies and proposals.
[1042] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1043] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1044] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1045] [Third embodiment]
[1046] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1047] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1048] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1049] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1050] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1051] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1052] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1053] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1054] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1055] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1056] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1057] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1058] The present invention is a system for facilitating communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. The following describes how the system of the present invention is specifically implemented.
[1059] Overall system configuration
[1060] The system consists of the following main components:
[1061] 1. Client device (augmented reality (AR) device worn by the user)
[1062] 2. Server (responsible for major processing on the cloud)
[1063] 3. User interface (smartphone app or web interface)
[1064] Initial setup and preparation phase
[1065] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[1066] Real-time support during business hours
[1067] The user puts on AR glasses or AR contact lenses, and the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. The server integrates the emotion estimation results with the current situation and generates optimal advice using a large-scale language model. This allows the user to receive suggestions for appropriate topics and actions in real time. The advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[1068] Follow-up
[1069] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user.
[1070] Specific examples
[1071] In the case of a date
[1072] 1. The day before the date
[1073] The user enters the other person's basic information into a smartphone app.
[1074] The server suggests a plan of "lunch at Cafe A" and "watch movie B."
[1075] The user selects "Watch Movie B" and the server reserves the tickets.
[1076] 2. The day of the date
[1077] The user puts on the AR glasses.
[1078] During the date, the AR glasses transmit video and audio to a server.
[1079] The server infers that the other person is "interested" based on their facial expression and tone of voice.
[1080] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[1081] The advice is displayed in the user's field of view, and the user provides the topic.
[1082] 3. After the date
[1083] The server generates a message saying, "Thank you for today. It was fun!"
[1084] The message is notified to the smartphone app, and the user confirms it before sending it.
[1085] In the case of sales
[1086] 1. Before the sales visit
[1087] The user enters the other company's information and the name of the person in charge into the smartphone app.
[1088] The server prepares initial proposals and topics based on past transaction history and related information.
[1089] 2. During a sales visit
[1090] The user puts on the AR contact lenses.
[1091] During business hours, the AR contact lenses capture video and audio and send them to a server.
[1092] The server analyzes the facial expressions and tone of voice of the person in charge and determines that they are "interested."
[1093] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[1094] The advice is displayed in the user's field of view and the user demonstrates it.
[1095] 3. After the visit
[1096] The server generates a follow-up email.
[1097] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[1098] Thus, the present invention is a system that supports users in real time in choosing the best topics and actions to take during dates or business trips, thereby improving the quality of communication.
[1099] The processing flow will be explained below.
[1100] Initial setup and preparation phase
[1101] Step 1:
[1102] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[1103] Step 2:
[1104] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received data, and retrieves matching data.
[1105] Step 3:
[1106] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[1107] Step 4:
[1108] The user checks the proposed date plans and sales strategies, selects one, and sends the selection information to the server.
[1109] Step 5:
[1110] The server makes the necessary reservations (restaurants, cinemas, etc.) based on the user's selection, and notifies the user's app that the reservation has been completed.
[1111] Real-time support during business hours
[1112] Step 1:
[1113] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[1114] Step 2:
[1115] The device uses its sensors to capture video and audio in real time, and sends the captured data to a server.
[1116] Step 3:
[1117] The server analyzes the received video data using an image recognition algorithm to infer emotions from facial expressions, and analyzes the audio data using a voice recognition algorithm to infer emotions from the tone and patterns of the voice.
[1118] Step 4:
[1119] The server sends the emotion estimation results to the user's smartphone app, which simultaneously analyzes the emotion information and the current dating or sales context using a large-scale language model to generate optimal advice.
[1120] Step 5:
[1121] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[1122] Follow-up
[1123] Step 1:
[1124] The user enters into the app that a date or business has ended, and the device sends all data up to the end to the server.
[1125] Step 2:
[1126] The server analyzes the data at the end, generates the next follow-up action to be taken, and notifies the smartphone app with the generated recommendation message.
[1127] Step 3:
[1128] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[1129] Example 1
[1130] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1131] In modern society, it is difficult to grasp appropriate topics and actions in real time in communication situations such as dating or business. In particular, accurately understanding one's own emotions and the emotions of others and acting accordingly is extremely important, but not easy to achieve. Conventional systems lack the ability to assess the other person's emotions and provide appropriate advice, resulting in a decline in the quality of communication. New technology is needed to solve these problems.
[1132] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1133] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for acquiring video and audio in real time and transmitting them to the server, and means for analyzing all data up to the end and generating the next action. This allows the user to receive suggestions for appropriate topics and actions in real time, enabling high-quality communication even during dates or business trips.
[1134] "Image recognition" is a technology that analyzes image data acquired using cameras and sensors to identify objects, people, features, etc.
[1135] "Emotion estimation" is a technology that determines the other person's psychological state and emotions by analyzing image and audio data.
[1136] "Speech recognition" is a technology that analyzes voice data acquired using devices such as microphones and identifies what a person is saying and the characteristics of their voice.
[1137] A "large-scale language model" is an artificial intelligence model that learns from large amounts of text data to generate and understand natural language.
[1138] An "augmented reality (AR) device" is a device that displays digital information superimposed on real-world visual information.
[1139] "Real-time" means processing acquired data immediately without delay and providing results instantly.
[1140] "Capturing video and audio" means collecting visual and audio information from the surroundings using devices such as cameras and microphones.
[1141] A "server" is a central computer system that provides services over a network.
[1142] "Analyzing all data" means extracting and analyzing information for a specific purpose from all acquired data.
[1143] "Generating the next action" means proposing the next action or response that the user should take based on the analysis results.
[1144] This invention is a system for facilitating smooth communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. Below, we will explain how to specifically implement the system of this invention.
[1145] Hardware and software used
[1146] Client device: An augmented reality (AR) device worn by a user (e.g., AR glasses, AR contact lenses).
[1147] Server: A computer system that handles major processing on the cloud.
[1148] User interface: Smartphone app and web interface.
[1149] Initial setup and preparation phase
[1150] 1. Enter information
[1151] The user starts the smartphone app and enters the other person's basic information (name, age, hobbies, previous conversations), which is then received by the server and stored in a database.
[1152] 2. Information Reception and Analysis
[1153] The server analyzes the received information and generates initial topics and suggestions based on past conversation history data. The server then uses a generative AI model to propose date plans and sales strategies.
[1154] 3. Plan proposal
[1155] The server proposes specific plans to the user, such as "lunch at Cafe A" or "watch movie B." The user selects one of the proposed plans and sends it to the server, which then arranges the reservations required for the selected plan.
[1156] Example prompt sentence:
[1157] Based on the other person's basic information, propose a date plan (e.g., lunch at Cafe A, watching Movie B).
[1158] Real-time support during business hours
[1159] 1. Wearing the device
[1160] Users wear AR glasses or AR contact lenses, and these devices capture video and audio in real time.
[1161] 2. Data Acquisition and Transmission
[1162] The device sends the captured video and audio to a server, which analyzes them and uses image and voice recognition technology to estimate the other person's emotions.
[1163] 3. Emotion Estimation and Advice Generation
[1164] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[1165] 4. Displaying Advice
[1166] The server sends the generated advice to the user's device, which then displays it in the user's field of view. For example, if the person speaking is smiling, the advice displayed might be, "Try talking about a movie you recently saw."
[1167] Example prompt sentence:
[1168] Based on the video and audio captured in real time, estimate the other person's emotions (e.g., interest) and generate appropriate topics of conversation.
[1169] Follow-up
[1170] 1. Notice of Termination
[1171] The user notifies the smartphone app that a date or business meeting has ended. The server collects all data up to that point and begins analyzing it.
[1172] 2. Suggested next actions
[1173] The server generates the next action to be taken. For example, the server generates a follow-up message such as "Thank you for today. I had a great time!" and notifies the user's smartphone app. The follow-up message is sent when the user confirms and sends it.
[1174] Example prompt sentence:
[1175] Generate a follow-up message to send after the date ends (e.g., "Thank you for today. I had a great time!").
[1176] In this way, the system of the present invention can support users in real time to choose the most appropriate topic and action in any situation, thereby realizing high-quality communication even during dates or business meetings.
[1177] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1178] Step 1:
[1179] Enter information
[1180] The user launches the smartphone app and enters the other person's basic information (name, age, hobbies, and previous conversation details).
[1181] Input: The user enters the following information into the input form: "Name = Tanaka Taro, Age = 28, Hobby = Reading, Previous conversation content = We talked about movies."
[1182] Output: The data entered by the user is saved in the app and sent to the server.
[1183] Step 2:
[1184] Information reception and analysis
[1185] The server receives the basic information sent by the user.
[1186] The server analyzes the received data and stores it in a database. By linking it with past conversation history data, it generates optimal initial topics and suggestions.
[1187] Input: Basic information data submitted by the user.
[1188] Data processing: Obtain past conversation data from the database and analyze the other person's interests.
[1189] Output: Generate a list of initial topics and suggestions.
[1190] Step 3:
[1191] Plan proposal
[1192] The server uses the generated AI model to propose specific date plans and sales strategies.
[1193] Example: Generate a plan that includes "Lunch at Cafe X" and "Watch Movie Y."
[1194] Input: A candidate list and a prompt sentence for the generative AI model.
[1195] Data processing: The AI model generates a date plan based on the prompt.
[1196] Output: A list of specific plans proposed to the user.
[1197] Step 4:
[1198] Plan selection and reservation arrangements
[1199] The user selects a plan suggested by the smartphone app. For example, they select "Watch Movie Y."
[1200] The server then arranges the necessary reservations for the selected plan, in this case booking movie tickets online.
[1201] Input: The user's selected plan.
[1202] Data processing: Connect to the online reservation system based on the selected plan and complete the reservation.
[1203] Output: Reservation confirmation information is sent from the server and notified to the user.
[1204] Step 5:
[1205] Wearing the device
[1206] The user puts on AR glasses or AR contact lenses before starting a date or business trip.
[1207] Input: The act of a user putting on a device.
[1208] Output: The AR device is initialized and begins capturing video and audio in real time.
[1209] Step 6:
[1210] Acquiring and Sending Data
[1211] The device captures video and audio in real time and transmits them to the server.
[1212] Input: Surrounding video and audio information.
[1213] Data processing: The device acquires the information, encodes it, and uploads it to the server.
[1214] Output: The data stream that the server receives in real time.
[1215] Step 7:
[1216] Emotion estimation and advice generation
[1217] The server analyzes the received video and audio and uses image and voice recognition technology to estimate the other person's emotions. For example, it detects the other person's smile and infers that they are "interested."
[1218] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[1219] Input: Received video and audio data.
[1220] Data processing: Analyzes video and audio to recognize the other person's emotions. A large-scale language model generates dialogue advice based on the recognition results.
[1221] Output: The generated advice is displayed in the user's field of view.
[1222] Step 8:
[1223] Displaying Advice
[1224] The server transmits the generated advice to the user's terminal.
[1225] The device displays advice in the user's field of view, such as "Tell me about a movie you recently saw."
[1226] Input: The generated advice.
[1227] Data processing: Converting advice into a visually displayable format.
[1228] Output: An advisory message that is displayed in the user's field of view.
[1229] Step 9:
[1230] Notice of Termination
[1231] The smartphone app notifies the user when a date or business meeting has ended.
[1232] Input: User clicks the "Exit" button.
[1233] Output: A completion notice is sent to the server.
[1234] Step 10:
[1235] Data analysis
[1236] The server collects all data up to the end point and begins analyzing it, for example, analyzing the content of the conversation that day and the changes in the other person's emotions.
[1237] Input: Termination notice and all collected data.
[1238] Data processing: Analyze conversation content and emotional data to extract trends.
[1239] Output: A report of the analysis results.
[1240] Step 11:
[1241] Suggested next actions
[1242] The server generates the next action to be taken, for example, a follow-up message such as "Thank you for today. I had a great time!"
[1243] Input: Analysis results.
[1244] Data processing: The next action is generated using a generative AI model.
[1245] Output: The actions and messages generated.
[1246] Step 12:
[1247] Notification and confirmation
[1248] The server generates a message and sends it to the user's smartphone app. The user confirms it and sends the message.
[1249] Input: The generated message.
[1250] Data processing: Notify the message and ask for confirmation.
[1251] Output: The message that is confirmed by the user and sent to the other party.
[1252] (Application example 1)
[1253] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1254] Existing customer service support systems make it difficult for staff to provide optimal product recommendations and customer service methods for each individual customer in a physical store. This makes it difficult to provide effective customer service to increase customer satisfaction, and it is difficult to increase store sales. In addition, because the system relies on the customer service skills of the staff, it is difficult for new or inexperienced staff to provide effective customer service.
[1255] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1256] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for analyzing the customer's facial expressions and acquiring facial emotion data, means for analyzing voice data and converting it into text, means for generating appropriate product proposals by combining the text data and the facial emotion data, and means for synthesizing the generated product proposals and notifying the staff. This enables the staff to grasp the customer's emotions and conversation content in real time and immediately receive optimal product proposals and customer service methods.
[1257] "Image recognition" is a technology that detects and analyzes the characteristics of specific objects or people from image data.
[1258] "Emotion estimation" means estimating a person's emotional state (e.g., joy, anger, sadness, etc.) based on image and audio data.
[1259] "Speech recognition" is a technology that analyzes voice data and converts it into text data.
[1260] A "large-scale language model" is an artificial intelligence model trained on vast amounts of text data that is capable of understanding and generating natural language.
[1261] An "augmented reality (AR) device" is a device that displays computer graphics overlaid on images of the real world.
[1262] "Analyzing facial expressions to obtain facial emotion data" means using image recognition technology to extract emotional information from a person's facial expressions and obtain that data.
[1263] "Analyzing voice data and converting it into text" means converting voice data into text data using voice recognition technology.
[1264] "Generating appropriate product proposals by combining text data and facial emotion data" means automatically generating optimal product proposals for customers using text data and facial emotion data.
[1265] "Synthesize voice and notify staff" means synthesizing the generated product suggestions and advice into voice and conveying it to staff.
[1266] This invention provides a system that makes full use of image recognition, voice recognition, and large-scale language models to support customer service in brick-and-mortar stores. Specific embodiments for implementing this system are described below.
[1267] System configuration
[1268] The system consists of the following main components:
[1269] 1. Client devices (including smart glasses and head-mounted displays)
[1270] 2. Server (located on the cloud and responsible for main processing)
[1271] 3. User interface (smartphone app or web interface)
[1272] Program Overview
[1273] Emotion estimation using image recognition
[1274] The client device captures the customer's facial expressions in real time while serving customers in a physical store and sends the image data to a server. The server then uses image recognition technology such as Amazon Rekognition to analyze the customer's emotions (e.g., joy, anger, sadness, etc.) and obtains facial emotion data.
[1275] Emotion estimation using speech recognition
[1276] The client device records the conversation with the customer and sends the audio data to the server. The server then converts the audio data into text data using voice recognition technology such as the Google Speech-to-Text API. This text data is then used to estimate emotions based on the conversation content.
[1277] Generating optimal suggestions using large-scale language models
[1278] The server uses a large-scale language model (e.g., GPT-4) to generate optimal product suggestions and customer service methods based on text data and facial emotion data. The generated suggestions are input as prompt sentences as follows:
[1279] Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation: "What products are you interested in recently?"
[1280] Display suggestions and speech synthesis
[1281] The generated suggestions are notified to staff using voice synthesis technology (such as pyttsx3). The user (staff member) receives this information in real time through smart glasses or a head-mounted display attached to the client device and can take appropriate action.
[1282] Specific examples
[1283] Example 1: Proposing a new product
[1284] When a customer asks a staff member, "What products do you recommend?", the system converts the voice data into text data and analyzes facial expressions to infer that the customer is "interested." Based on this information, a large-scale language model generates a suggestion such as "Recommend a newly arrived product," and notifies the staff member via voice synthesis.
[1285] Example 2: Follow-up suggestions
[1286] After the staff member has finished serving the customer, the system analyzes the acquired data and generates an appropriate follow-up action (e.g., sending a thank-you message) and notifies the user interface. The user can then confirm the suggestion and take action if necessary.
[1287] As a result, this system allows staff to understand customers' emotions and conversation content in real time, enabling them to instantly receive optimal product suggestions and customer service methods.
[1288] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1289] Step 1:
[1290] The user wears the client terminal.
[1291] How it works: The customer wears smart glasses or a head-mounted display, which allows the device's camera and microphone to capture the customer's video and audio in real time.
[1292] Input: The client device worn by the user.
[1293] Output: Real-time video and audio capture.
[1294] Step 2:
[1295] The terminal captures the customer's video and audio and sends them to the server.
[1296] How it works: The device's camera captures the customer's facial expressions, and the microphone records the conversation. This data is then sent to a cloud server.
[1297] Input: Video and audio data acquired by the device.
[1298] Output: Video and audio data sent to the server.
[1299] Step 3:
[1300] The server uses image recognition technology to analyze the customer's facial expressions and obtain facial emotion data.
[1301] How it works: The server runs an image recognition algorithm such as Amazon Rekognition to estimate the customer's emotions from their facial expressions. For example, the results can be "happy" or "interested" based on the facial expression.
[1302] Input: Video data sent to the server.
[1303] Output: Facial emotion data (e.g., "HAPPY", confidence level 95.0%).
[1304] Step 4:
[1305] The server uses voice recognition technology to convert the voice data into text data.
[1306] What happens: The server runs a speech recognition algorithm, such as the Google Speech-to-Text API, to convert the conversation into text. For example, it might generate something like, "What products are you interested in these days?"
[1307] Input: The audio data sent to the server.
[1308] Output: Text data (e.g., "What products have you been interested in recently?").
[1309] Step 5:
[1310] The server combines text data and facial emotion data to input prompts into a large-scale language model to generate optimal product suggestions.
[1311] Specific operation: The server uses a large-scale language model (e.g., GPT-4) to generate a prompt sentence and then generates appropriate product suggestions based on it. An example of a prompt sentence is "Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation content: 'What products are you interested in recently?'"
[1312] Input: Text data and facial emotion data.
[1313] Output: Product suggestions (e.g. "Recommend new products").
[1314] Step 6:
[1315] The product suggestions generated by the server are notified to staff using voice synthesis technology.
[1316] Specific operation: The server uses a speech synthesis algorithm such as pyttsx3 to convert the generated product suggestions into voice and send it to the client terminal. The user (staff member) receives the suggestions by voice and can respond appropriately.
[1317] Input: Server-generated product suggestions.
[1318] Output: A voice-synthesized product suggestion notification.
[1319] Step 7:
[1320] Based on the user-generated proposals, appropriate products are proposed to the customer.
[1321] Specific operation: The user (staff member) uses the product proposal notified by voice from the client terminal as a reference to explain and propose new products to the customer.
[1322] Input: Speech-synthesized product suggestions.
[1323] Output: Product recommendations to customers.
[1324] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1325] This invention is a system that analyzes the emotions of users and their partners in real time during dates and business meetings, and suggests optimal topics and behaviors. This system achieves highly accurate emotion analysis by combining image recognition, speech recognition technology, large-scale language models, and an emotion engine.
[1326] Overall system configuration
[1327] The system consists of the following main components:
[1328] 1. Client device (augmented reality (AR) device worn by the user)
[1329] 2. Server (responsible for major processing on the cloud)
[1330] 3. Emotion engine (analyzing user and other party emotions in real time)
[1331] 4. User interface (smartphone app or web interface)
[1332] Initial setup and preparation phase
[1333] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[1334] Real-time support during business hours
[1335] When a user wears AR glasses or AR contact lenses, the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. At the same time, an emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[1336] The server integrates the emotional information of both the user and the other party, analyzes the current dating or sales context using a large-scale language model, and generates optimal advice. The generated advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[1337] Furthermore, the emotion engine learns from the user's past emotional data, analyzes their emotional tendencies, and provides advice in response to predicted emotional changes. This makes it easier for the user to understand their own emotional state and continue the conversation calmly.
[1338] Follow-up
[1339] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user. The notified message can be sent after the user confirms it.
[1340] Specific examples
[1341] In the case of a date
[1342] 1. The day before the date
[1343] The user enters the other person's basic information into a smartphone app.
[1344] The server suggests plans such as "lunch at Cafe A" or "watching Movie B."
[1345] The user selects "Watch Movie B" and the server reserves the tickets.
[1346] 2. The day of the date
[1347] The user puts on the AR glasses.
[1348] During the date, the AR glasses transmit video and audio to a server.
[1349] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[1350] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[1351] The advice is displayed in the user's field of view, and the user provides the topic.
[1352] 3. After the date
[1353] The server generates a message saying, "Thank you for today. It was fun!"
[1354] The message is notified to the smartphone app, and the user confirms it before sending it.
[1355] In the case of sales
[1356] 1. Before the sales visit
[1357] The user enters the other company's information and the name of the person in charge into the smartphone app.
[1358] The server prepares initial proposals and topics based on past transaction history and related information.
[1359] 2. During a sales visit
[1360] The user puts on the AR contact lenses.
[1361] During business hours, the AR contact lenses capture video and audio and send them to a server.
[1362] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[1363] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[1364] The advice is displayed in the user's field of view and the user demonstrates it.
[1365] 3. After the visit
[1366] The server generates a follow-up email.
[1367] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[1368] In this way, the present invention is a system that improves the quality of communication by analyzing the emotions of the user and the other party in real time and suggesting optimal topics and actions. The introduction of an emotion engine makes it easier to understand the user's own emotions, enabling more effective support.
[1369] The processing flow will be explained below.
[1370] Initial setup and preparation phase
[1371] Step 1:
[1372] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[1373] Step 2:
[1374] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received information, and retrieves matching data.
[1375] Step 3:
[1376] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[1377] Step 4:
[1378] The user checks the proposed date plans and sales strategies and selects one. The selected information is sent to the server.
[1379] Step 5:
[1380] The server makes the necessary reservations (for restaurants, cinemas, etc.) based on the user's selection, and notifies the user's smartphone app that the reservation has been completed.
[1381] Real-time support during business hours
[1382] Step 1:
[1383] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[1384] Step 2:
[1385] The device uses its sensors to capture video and audio in real time, and the captured data is sent to a server.
[1386] Step 3:
[1387] The server analyzes the video data it receives using an image recognition algorithm to infer emotions from facial expressions, and it also analyzes audio data using a voice recognition algorithm to infer emotions from the tone and patterns of voice.
[1388] Step 4:
[1389] The server creates the other person's emotional information based on the emotion estimation results. At the same time, the emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[1390] Step 5:
[1391] The server integrates the emotional information of both the user and the other party and generates optimal advice using a large-scale language model. The advice is then displayed in the user's field of view.
[1392] Step 6:
[1393] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[1394] Follow-up
[1395] Step 1:
[1396] The user enters into the app that a date or business trip has ended, and the device sends all data up to the end to the server.
[1397] Step 2:
[1398] The server analyzes the data at the end and generates a recommended action to be taken next, which is then sent to the user's smartphone app.
[1399] Step 3:
[1400] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[1401] Example 2
[1402] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1403] In conventional dating and sales situations, it has been difficult for users to accurately read the other person's emotions and decide on appropriate topics and actions. Furthermore, it has been difficult for users to manage their own emotions while having an effective conversation. This has led to a decline in the quality of communication and a decrease in the success rate of dates and sales.
[1404] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through voice recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and an emotion engine for analyzing the user's emotions in real time. This configuration enables the user to grasp the emotions of others and themselves in real time and take optimal actions and responses.
[1405] "Image recognition" is the technology that extracts information from digital images and videos and identifies specific patterns and features.
[1406] "Speech recognition" is a technology that analyzes voice data and converts spoken words into text data.
[1407] A "large-scale language model" is an artificial intelligence model trained on large amounts of text data, and is a model that excels at understanding and generating natural language.
[1408] An "augmented reality (AR) device" is a device that overlays digital information onto a user's real-world field of vision. Examples include AR glasses and AR contact lenses.
[1409] The "emotion engine" is a system that analyzes the emotional state of the user and the other person in real time and provides appropriate feedback and advice based on the results.
[1410] A "date plan" refers to a proposal or action plan for systematically preparing the progress of a date.
[1411] A "sales strategy" is a plan for effective contact methods and presentations in sales activities.
[1412] "Real-time data" is data that represents ongoing events or conditions and should be processed immediately.
[1413] "Follow-up actions" are subsequent actions that should be taken after a date or business meeting, with the purpose of maintaining and strengthening the relationship with the other person.
[1414] "Initial settings" refers to the setup work required before starting to use the system, in which basic information about the user and the other party is entered and preparations required for system operation are made.
[1415] The system based on this invention analyzes the emotions of the user and the other party in real time during dating or business, and suggests optimal topics and behaviors. This system includes the following main components: image recognition, speech recognition, a large-scale language model, an augmented reality (AR) device, and an emotion engine.
[1416] Hardware and software used
[1417] Hardware: Augmented reality (AR) glasses, augmented reality (AR) contact lenses, smartphones
[1418] Software: Image recognition software (e.g., Google Cloud Vision), speech recognition software (e.g., Google Speech-to-Text), large-scale language models (e.g., OpenAI GPT-3), sentiment analysis engines (e.g., Affectiva)
[1419] Initial Setup and Data Entry
[1420] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the basic information sent by the user, coordinates it with past conversation history data to perform initial setup, and generates optimal initial topics and suggestions (e.g., date plans or sales strategies). The user selects a suggested plan on the smartphone app, and the server arranges reservations and other preparations based on the selected plan.
[1421] Real-time data acquisition and analysis
[1422] The user wears an AR device (e.g., AR glasses or AR contact lenses), and the device captures video and audio in real time. The captured data is sent to a server, which analyzes it using image and voice recognition software. The analysis results are processed by an emotion engine, which analyzes the emotional state of the other person and the user themselves.
[1423] Generating and displaying advice
[1424] Based on the analysis results, the server uses a large-scale language model to generate optimal topics and behaviors. The generated advice is sent to the user's AR device and displayed as an overlay in the user's field of view. The user can refer to the displayed advice to proceed with the conversation.
[1425] Follow-up
[1426] After a date or business meeting ends, the user notifies the end of the event via a smartphone app. The server analyzes all data up to the end of the event and generates the next follow-up action (such as sending a message). The generated message is notified to the user, who can then send it after confirming it.
[1427] Specific examples
[1428] Specific examples of dating
[1429] 1. The day before the date
[1430] The user enters the other person's basic information into a smartphone app.
[1431] The server will suggest plans such as "lunch at a cafe" or "watching a movie."
[1432] The user selects "Watch a Movie" and the server reserves the tickets.
[1433] 2. The day of the date
[1434] The user puts on the AR glasses.
[1435] During the date, the AR glasses transmit video and audio to a server.
[1436] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[1437] A large-scale language model generates advice like, "Tell me about a movie you recently saw."
[1438] The advice is displayed in the user's field of view, and the user provides the topic.
[1439] 3. After the date
[1440] The server generates a message saying "Thanks for today, see you again soon!"
[1441] The message is notified to the smartphone app, and the user confirms it before sending it.
[1442] Specific examples of sales
[1443] 1. Before the sales visit
[1444] The user enters the other company's information and the name of the person in charge into the smartphone app.
[1445] The server references transaction history and related information to prepare initial proposals and topics.
[1446] 2. During a sales visit
[1447] The user puts on the AR contact lenses.
[1448] During business hours, the AR contact lenses capture video and audio and send it to a server.
[1449] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[1450] A large-scale language model generates the advice "Introduce and demonstrate the new product."
[1451] The advice is displayed in the user's field of view and the user demonstrates it.
[1452] 3. After business hours
[1453] The server generates a follow-up email to notify the user.
[1454] The user checks the email and then sends it.
[1455] Prompt Sentence Examples
[1456] "If the other person seems to be having fun, what topic should I bring up next?"
[1457] "What should I do when my partner seems a little cranky?"
[1458] "Suggest topics based on topics that your partner has shown interest in on previous dates."
[1459] In this way, the system based on this invention analyzes the emotions of the user and the other party in real time and suggests optimal topics and actions, thereby improving the quality of communication. The introduction of an emotion engine makes it easier for users to understand their own emotions, enabling more effective support.
[1460] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1461] Step 1:
[1462] The user launches the smartphone app. The user enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives this basic information and performs initial setup in conjunction with past conversation history data. Specifically, the server stores the entered name and hobbies of the other person in a database and compares them with past conversation history data. As an output, the most suitable initial topic and suggestions for the other person are generated based on the comparison results.
[1463] Step 2:
[1464] The user selects a proposed plan (e.g., a date plan or a sales strategy) on a smartphone app. The server receives the selected plan and arranges and prepares the necessary reservations based on that plan. At this time, the server connects to the reservation system and secures the necessary resources (e.g., movie tickets, restaurant reservations). As an output, the confirmed date plan or sales strategy is generated and notified to the user.
[1465] Step 3:
[1466] A user puts on an augmented reality (AR) device (e.g., AR glasses or AR contact lenses). The device captures video and audio in real time. This data is collected through the device's camera and microphone. The input real-time data is sent to a server, which prepares it for data analysis. Specifically, video data is sent to the server in streaming format, and audio data is transferred in real time. Data begins to be sent to the server as output.
[1467] Step 4:
[1468] The server analyzes the video data using image recognition software (e.g., Google Cloud Vision) and the audio data using voice recognition software (e.g., Google Speech-to-Text). In this analysis, the other person's facial expressions and tone of voice are extracted as important elements. The input is video and audio data transmitted in real time, and the output is data indicating the emotional state. Specifically, the server extracts facial expression features from the video data and extracts tone of voice and word features from the audio data.
[1469] Step 5:
[1470] The emotion engine analyzes the emotional state of the user and the other party in real time based on the analysis results. The input is the analysis results of image recognition and voice recognition, and the output is emotion analysis data. The emotion engine integrates this data and grasps the psychological state of the user and the other party in real time. Specifically, the emotion engine also monitors changes in the user's heart rate and breathing.
[1471] Step 6:
[1472] The server inputs the sentiment analysis data into a large-scale language model (e.g., OpenAI GPT-3) to generate optimal topics and behaviors. The input is the sentiment analysis data and provided context information, and the output is specific advice. By sending prompts to the large-scale language model, responses are generated to questions such as, "If the other person seems to be having fun, what topic should I bring up next?"
[1473] Step 7:
[1474] The server compiles the generated advice and sends it to the user's AR device. The input is advice generated by a large-scale language model, and the output is information overlaid on the user's field of view. Specifically, the server sends appropriate overlay information to the AR device, and the advice is displayed in the user's field of view.
[1475] Step 8:
[1476] The user proceeds with the dialogue while referring to the displayed advice. The input is the advice displayed on the AR device, and the output is the user's actions. The user interacts with the other party based on the advice, and the results are sent to the server as data acquired in real time. Specifically, the dialogue content and the other party's reactions are continuously acquired in real time.
[1477] Step 9:
[1478] After a date or business meeting ends, the user notifies the smartphone app that the event is over. The server analyzes all data up to the end and generates the follow-up action to be taken next. The input is all data up to the end, and the output is the follow-up action (e.g., sending a message). In concrete terms, the server comprehensively analyzes all the data obtained and determines the next step.
[1479] Step 10:
[1480] The server notifies the user of a follow-up message generated by the server. The user confirms and sends the message. The input is the message generated by the server, and the output is the message confirmed and sent by the user. Specifically, the user confirms the message in the smartphone app, edits or corrects it as necessary, and then taps the send button.
[1481] (Application example 2)
[1482] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1483] In conventional marketing activities, it has been difficult to instantly grasp the emotional reactions of clients and customers and propose optimal advertising strategies in real time. This has led to problems such as not being able to maximize the effectiveness of advertising and insufficient communication with clients. In particular, it has been difficult to provide effective feedback in situations where it is necessary to respond quickly to client reactions during a conversation.
[1484] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1485] In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through speech recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and a means for analyzing the emotional responses of conversation partners in real time during marketing activities and generating optimal advertising strategies and proposals. This makes it possible to analyze the emotions of clients or customers in real time and quickly generate optimal advertising strategies and proposals accordingly. This system is effective for marketers to receive instant feedback during conversations and promote effective communication with clients.
[1486] "Image recognition" is a technology that extracts and analyzes specific information and features from image data acquired using devices such as cameras.
[1487] "Means for estimating emotions" refers to technology that uses image recognition and voice recognition to analyze a person's facial expressions and tone of voice, and then estimates emotions based on the results.
[1488] "Speech recognition" is a technology that analyzes voice data acquired through devices such as microphones and extracts characteristics of words and voices.
[1489] A "large-scale language model" is a machine learning model trained using massive amounts of text data, and has the ability to understand and generate natural language.
[1490] "Means for generating optimal topics and behaviors" refers to a technology that uses a large-scale language model to generate suggestions for optimal topics and actions according to the situation.
[1491] A "user-worn augmented reality (AR) device" is a device worn by a user that can overlay digital information on the user's field of vision.
[1492] "Marketing activities" refers to various sales activities aimed at promoting products and services.
[1493] "Means for analyzing the emotional responses of a conversation partner in real time" refers to technology that instantly analyzes the emotions and responses of the other person during a conversation and provides that information.
[1494] An "advertising strategy" refers to a plan or method for effectively promoting a particular product or service.
[1495] The "means for generating suggestions" is a technology that suggests optimal actions and conversation content to users based on analyzed data.
[1496] The present invention provides a system for analyzing the emotions of a conversation partner in real time during marketing activities and for optimizing advertising strategies and proposals. The system is realized using image recognition, speech recognition, a large-scale language model, a wearable augmented reality (AR) device, and an emotion analysis engine. Specific embodiments for implementing the present invention are described below.
[1497] Overall system configuration
[1498] The system consists of the following main components:
[1499] 1. Client device (augmented reality (AR) device worn by the user)
[1500] 2. Server (responsible for major processing on the cloud)
[1501] 3. Emotion engine (analyzing the emotions of the person interacting with you and the user in real time)
[1502] 4. User interface (smartphone app or web interface)
[1503] Initial setup and preparation phase
[1504] First, the user launches the smartphone app and registers the other party's basic information (such as name, age, hobbies, and response to previous advertisements) in natural language. The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate an optimal initial advertising strategy and proposal for the other party and propose a marketing plan. The user selects the proposed plan and sends it to the server, completing preparations for marketing activities.
[1505] Real-time support during marketing activities
[1506] The user wears an augmented reality (AR) device, which captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the emotion of the person they are interacting with. At the same time, an emotion engine analyzes and monitors the user's emotional state in real time.
[1507] The server integrates the emotional information of the user and the conversation partner, analyzes the context of current marketing activities using a large-scale language model, and generates optimal advertising strategies and proposals. The generated proposals are displayed in the user's field of view, and the user continues the dialogue using them as reference. Furthermore, the emotion engine learns the user's past emotional data, analyzes their emotional tendencies, and makes proposals that correspond to predicted emotional changes.
[1508] Follow-up
[1509] After the marketing activity is completed, the user notifies the smartphone app. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending an email), and notifies the user. The notified message can be sent after the user confirms it.
[1510] Hardware and software used
[1511] Hardware:
[1512] Smartphone
[1513] camera
[1514] Augmented Reality (AR) Devices
[1515] software:
[1516] OpenCV: Image Recognition Library
[1517] dlib: face detection library
[1518] DeepFace: A library for emotion analysis
[1519] OpenAI API: Large-scale language models
[1520] Specific examples
[1521] During a meeting with an important client, a marketing manager at an advertising agency uses a smartphone camera to analyze the client's facial expressions. The manager and his colleagues then review the emotion analysis results to devise strategies and make optimal proposals in real time.
[1522] Prompt Sentence Examples
[1523] My client is expressing "delight" emotions. What advertising strategy should I suggest next?
[1524] My client is expressing "anger" emotions. How do I fix this next?
[1525] My client is expressing feelings of sadness. How should I follow up next?
[1526] As described above, the present invention provides a system that combines sentiment analysis and large-scale language models to provide optimal advertising strategies and proposals in real time for marketing activities.
[1527] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1528] Step 1:
[1529] The server receives basic information about the other party (such as name, age, hobbies, and response to previous advertisements) from the user's device (smartphone app). This allows the necessary information to be accumulated on the server. Natural language data entered on the user's device is sent to the server, where it is analyzed and saved.
[1530] Step 2:
[1531] The server references the received contact's basic information and past conversation history data to perform initial setup. It also generates and proposes an appropriate initial advertising strategy and marketing plan. The plan is generated using a large-scale language model, using the contact's basic information and past conversation history data as input.
[1532] Step 3:
[1533] The user selects a proposed plan and sends it to the server. The server determines a marketing activity plan based on the selected plan and notifies the user's device. The selected plan and related data are processed and stored on the server.
[1534] Step 4:
[1535] A user wears an augmented reality (AR) device and starts interacting with the client. The AR device captures video and audio in real time and sends them to the server. The video and audio data are sent to the server as input.
[1536] Step 5:
[1537] The server analyzes the received video data using OpenCV and dlib to perform facial recognition and emotion estimation (image recognition). At the same time, it analyzes the received audio data using DeepFace to estimate emotions (speech recognition). The results of the emotion analysis are generated and saved on the server.
[1538] Step 6:
[1539] The emotion analysis engine monitors the emotional state of the user and the person they are interacting with based on the analyzed emotion data. The monitoring data is created on the server.
[1540] Step 7:
[1541] It uses a large-scale language model to analyze the context of current marketing activities and generate optimal advertising strategies and proposals. Using prompt sentences, it generates specific advertising strategies and proposals and displays them in the user's field of view. The generated advertising strategies and proposals are displayed on the user's AR device.
[1542] Step 8:
[1543] The user visually confirms the suggestions through the AR device and applies them during the dialogue to communicate with the client. The user advances the conversation based on the generated strategies and suggestions.
[1544] Step 9:
[1545] After completing a marketing activity, the user notifies the server using the smartphone app.
[1546] Step 10:
[1547] The server comprehensively analyzes all log data and sentiment analysis data and generates the next follow-up action to be taken. The generated action and the content of the follow-up are notified to the user's device.
[1548] Step 11:
[1549] The user confirms the follow-up action (e.g., sending an email) and sends it as is if necessary. The follow-up action is executed from the user's device.
[1550] Through the above steps, the system of the present invention can analyze the emotions of the person in conversation in real time during marketing activities and make optimal advertising strategies and proposals.
[1551] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1552] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1553] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1554] [Fourth embodiment]
[1555] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1556] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1557] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1558] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1559] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1560] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1561] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1562] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1563] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1564] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1565] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1566] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1567] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1568] The present invention is a system for facilitating communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. The following describes how the system of the present invention is specifically implemented.
[1569] Overall system configuration
[1570] The system consists of the following main components:
[1571] 1. Client device (augmented reality (AR) device worn by the user)
[1572] 2. Server (responsible for major processing on the cloud)
[1573] 3. User interface (smartphone app or web interface)
[1574] Initial setup and preparation phase
[1575] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[1576] Real-time support during business hours
[1577] The user puts on AR glasses or AR contact lenses, and the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. The server integrates the emotion estimation results with the current situation and generates optimal advice using a large-scale language model. This allows the user to receive suggestions for appropriate topics and actions in real time. The advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[1578] Follow-up
[1579] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user.
[1580] Specific examples
[1581] In the case of a date
[1582] 1. The day before the date
[1583] The user enters the other person's basic information into a smartphone app.
[1584] The server suggests a plan of "lunch at Cafe A" and "watch movie B."
[1585] The user selects "Watch Movie B" and the server reserves the tickets.
[1586] 2. The day of the date
[1587] The user puts on the AR glasses.
[1588] During the date, the AR glasses transmit video and audio to a server.
[1589] The server infers that the other person is "interested" based on their facial expression and tone of voice.
[1590] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[1591] The advice is displayed in the user's field of view, and the user provides the topic.
[1592] 3. After the date
[1593] The server generates a message saying, "Thank you for today. It was fun!"
[1594] The message is notified to the smartphone app, and the user confirms it before sending it.
[1595] In the case of sales
[1596] 1. Before the sales visit
[1597] The user enters the other company's information and the name of the person in charge into the smartphone app.
[1598] The server prepares initial proposals and topics based on past transaction history and related information.
[1599] 2. During a sales visit
[1600] The user puts on the AR contact lenses.
[1601] During business hours, the AR contact lenses capture video and audio and send them to a server.
[1602] The server analyzes the facial expressions and tone of voice of the person in charge and determines that they are "interested."
[1603] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[1604] The advice is displayed in the user's field of view and the user demonstrates it.
[1605] 3. After the visit
[1606] The server generates a follow-up email.
[1607] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[1608] Thus, the present invention is a system that supports users in real time in choosing the best topics and actions to take during dates or business trips, thereby improving the quality of communication.
[1609] The processing flow will be explained below.
[1610] Initial setup and preparation phase
[1611] Step 1:
[1612] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[1613] Step 2:
[1614] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received data, and retrieves matching data.
[1615] Step 3:
[1616] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[1617] Step 4:
[1618] The user checks the proposed date plans and sales strategies, selects one, and sends the selection information to the server.
[1619] Step 5:
[1620] The server makes the necessary reservations (restaurants, cinemas, etc.) based on the user's selection, and notifies the user's app that the reservation has been completed.
[1621] Real-time support during business hours
[1622] Step 1:
[1623] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[1624] Step 2:
[1625] The device uses its sensors to capture video and audio in real time, and sends the captured data to a server.
[1626] Step 3:
[1627] The server analyzes the received video data using an image recognition algorithm to infer emotions from facial expressions, and analyzes the audio data using a voice recognition algorithm to infer emotions from the tone and patterns of the voice.
[1628] Step 4:
[1629] The server sends the emotion estimation results to the user's smartphone app, which simultaneously analyzes the emotion information and the current dating or sales context using a large-scale language model to generate optimal advice.
[1630] Step 5:
[1631] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[1632] Follow-up
[1633] Step 1:
[1634] The user enters into the app that a date or business has ended, and the device sends all data up to the end to the server.
[1635] Step 2:
[1636] The server analyzes the data at the end, generates the next follow-up action to be taken, and notifies the smartphone app with the generated recommendation message.
[1637] Step 3:
[1638] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[1639] Example 1
[1640] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1641] In modern society, it is difficult to grasp appropriate topics and actions in real time in communication situations such as dating or business. In particular, accurately understanding one's own emotions and the emotions of others and acting accordingly is extremely important, but not easy to achieve. Conventional systems lack the ability to assess the other person's emotions and provide appropriate advice, resulting in a decline in the quality of communication. New technology is needed to solve these problems.
[1642] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1643] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for acquiring video and audio in real time and transmitting them to the server, and means for analyzing all data up to the end and generating the next action. This allows the user to receive suggestions for appropriate topics and actions in real time, enabling high-quality communication even during dates or business trips.
[1644] "Image recognition" is a technology that analyzes image data acquired using cameras and sensors to identify objects, people, features, etc.
[1645] "Emotion estimation" is a technology that determines the other person's psychological state and emotions by analyzing image and audio data.
[1646] "Speech recognition" is a technology that analyzes voice data acquired using devices such as microphones and identifies what a person is saying and the characteristics of their voice.
[1647] A "large-scale language model" is an artificial intelligence model that learns from large amounts of text data to generate and understand natural language.
[1648] An "augmented reality (AR) device" is a device that displays digital information superimposed on real-world visual information.
[1649] "Real-time" means processing acquired data immediately without delay and providing results instantly.
[1650] "Capturing video and audio" means collecting visual and audio information from the surroundings using devices such as cameras and microphones.
[1651] A "server" is a central computer system that provides services over a network.
[1652] "Analyzing all data" means extracting and analyzing information for a specific purpose from all acquired data.
[1653] "Generating the next action" means proposing the next action or response that the user should take based on the analysis results.
[1654] This invention is a system for facilitating smooth communication in dating and business situations. This system utilizes image recognition, speech recognition, and large-scale language models to suggest appropriate topics and behaviors in real time. Below, we will explain how to specifically implement the system of this invention.
[1655] Hardware and software used
[1656] Client device: An augmented reality (AR) device worn by a user (e.g., AR glasses, AR contact lenses).
[1657] Server: A computer system that handles major processing on the cloud.
[1658] User interface: Smartphone app and web interface.
[1659] Initial setup and preparation phase
[1660] 1. Enter information
[1661] The user starts the smartphone app and enters the other person's basic information (name, age, hobbies, previous conversations), which is then received by the server and stored in a database.
[1662] 2. Information Reception and Analysis
[1663] The server analyzes the received information and generates initial topics and suggestions based on past conversation history data. The server then uses a generative AI model to propose date plans and sales strategies.
[1664] 3. Plan proposal
[1665] The server proposes specific plans to the user, such as "lunch at Cafe A" or "watch movie B." The user selects one of the proposed plans and sends it to the server, which then arranges the reservations required for the selected plan.
[1666] Example prompt sentence:
[1667] Based on the other person's basic information, propose a date plan (e.g., lunch at Cafe A, watching Movie B).
[1668] Real-time support during business hours
[1669] 1. Wearing the device
[1670] Users wear AR glasses or AR contact lenses, and these devices capture video and audio in real time.
[1671] 2. Data Acquisition and Transmission
[1672] The device sends the captured video and audio to a server, which analyzes them and uses image and voice recognition technology to estimate the other person's emotions.
[1673] 3. Emotion Estimation and Advice Generation
[1674] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[1675] 4. Displaying Advice
[1676] The server sends the generated advice to the user's device, which then displays it in the user's field of view. For example, if the person speaking is smiling, the advice displayed might be, "Try talking about a movie you recently saw."
[1677] Example prompt sentence:
[1678] Based on the video and audio captured in real time, estimate the other person's emotions (e.g., interest) and generate appropriate topics of conversation.
[1679] Follow-up
[1680] 1. Notice of Termination
[1681] The user notifies the smartphone app that a date or business meeting has ended. The server collects all data up to that point and begins analyzing it.
[1682] 2. Suggested next actions
[1683] The server generates the next action to be taken. For example, the server generates a follow-up message such as "Thank you for today. I had a great time!" and notifies the user's smartphone app. The follow-up message is sent when the user confirms and sends it.
[1684] Example prompt sentence:
[1685] Generate a follow-up message to send after the date ends (e.g., "Thank you for today. I had a great time!").
[1686] In this way, the system of the present invention can support users in real time to choose the most appropriate topic and action in any situation, thereby realizing high-quality communication even during dates or business meetings.
[1687] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1688] Step 1:
[1689] Enter information
[1690] The user launches the smartphone app and enters the other person's basic information (name, age, hobbies, and previous conversation details).
[1691] Input: The user enters the following information into the input form: "Name = Tanaka Taro, Age = 28, Hobby = Reading, Previous conversation content = We talked about movies."
[1692] Output: The data entered by the user is saved in the app and sent to the server.
[1693] Step 2:
[1694] Information reception and analysis
[1695] The server receives the basic information sent by the user.
[1696] The server analyzes the received data and stores it in a database. By linking it with past conversation history data, it generates optimal initial topics and suggestions.
[1697] Input: Basic information data submitted by the user.
[1698] Data processing: Obtain past conversation data from the database and analyze the other person's interests.
[1699] Output: Generate a list of initial topics and suggestions.
[1700] Step 3:
[1701] Plan proposal
[1702] The server uses the generated AI model to propose specific date plans and sales strategies.
[1703] Example: Generate a plan that includes "Lunch at Cafe X" and "Watch Movie Y."
[1704] Input: A candidate list and a prompt sentence for the generative AI model.
[1705] Data processing: The AI model generates a date plan based on the prompt.
[1706] Output: A list of specific plans proposed to the user.
[1707] Step 4:
[1708] Plan selection and reservation arrangements
[1709] The user selects a plan suggested by the smartphone app. For example, they select "Watch Movie Y."
[1710] The server then arranges the necessary reservations for the selected plan, in this case booking movie tickets online.
[1711] Input: The user's selected plan.
[1712] Data processing: Connect to the online reservation system based on the selected plan and complete the reservation.
[1713] Output: Reservation confirmation information is sent from the server and notified to the user.
[1714] Step 5:
[1715] Wearing the device
[1716] The user puts on AR glasses or AR contact lenses before starting a date or business trip.
[1717] Input: The act of a user putting on a device.
[1718] Output: The AR device is initialized and begins capturing video and audio in real time.
[1719] Step 6:
[1720] Acquiring and Sending Data
[1721] The device captures video and audio in real time and transmits them to the server.
[1722] Input: Surrounding video and audio information.
[1723] Data processing: The device acquires the information, encodes it, and uploads it to the server.
[1724] Output: The data stream that the server receives in real time.
[1725] Step 7:
[1726] Emotion estimation and advice generation
[1727] The server analyzes the received video and audio and uses image and voice recognition technology to estimate the other person's emotions. For example, it detects the other person's smile and infers that they are "interested."
[1728] The server integrates the emotion estimation results with the current situation and generates appropriate advice using a large-scale language model.
[1729] Input: Received video and audio data.
[1730] Data processing: Analyzes video and audio to recognize the other person's emotions. A large-scale language model generates dialogue advice based on the recognition results.
[1731] Output: The generated advice is displayed in the user's field of view.
[1732] Step 8:
[1733] Displaying Advice
[1734] The server transmits the generated advice to the user's terminal.
[1735] The device displays advice in the user's field of view, such as "Tell me about a movie you recently saw."
[1736] Input: The generated advice.
[1737] Data processing: Converting advice into a visually displayable format.
[1738] Output: An advisory message that is displayed in the user's field of view.
[1739] Step 9:
[1740] Notice of Termination
[1741] The smartphone app notifies the user when a date or business meeting has ended.
[1742] Input: User clicks the "Exit" button.
[1743] Output: A completion notice is sent to the server.
[1744] Step 10:
[1745] Data analysis
[1746] The server collects all data up to the end point and begins analyzing it, for example, analyzing the content of the conversation that day and the changes in the other person's emotions.
[1747] Input: Termination notice and all collected data.
[1748] Data processing: Analyze conversation content and emotional data to extract trends.
[1749] Output: A report of the analysis results.
[1750] Step 11:
[1751] Suggested next actions
[1752] The server generates the next action to be taken, for example, a follow-up message such as "Thank you for today. I had a great time!"
[1753] Input: Analysis results.
[1754] Data processing: The next action is generated using a generative AI model.
[1755] Output: The actions and messages generated.
[1756] Step 12:
[1757] Notification and confirmation
[1758] The server generates a message and sends it to the user's smartphone app. The user confirms it and sends the message.
[1759] Input: The generated message.
[1760] Data processing: Notify the message and ask for confirmation.
[1761] Output: The message that is confirmed by the user and sent to the other party.
[1762] (Application example 1)
[1763] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1764] Existing customer service support systems make it difficult for staff to provide optimal product recommendations and customer service methods for each individual customer in a physical store. This makes it difficult to provide effective customer service to increase customer satisfaction, and it is difficult to increase store sales. In addition, because the system relies on the customer service skills of the staff, it is difficult for new or inexperienced staff to provide effective customer service.
[1765] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1766] In this invention, the server includes means for estimating emotions through image recognition, means for estimating emotions through voice recognition, means for generating optimal topics and behaviors using a large-scale language model, means for displaying information on an augmented reality (AR) device worn by the user, means for analyzing the customer's facial expressions and acquiring facial emotion data, means for analyzing voice data and converting it into text, means for generating appropriate product proposals by combining the text data and the facial emotion data, and means for synthesizing the generated product proposals and notifying the staff. This enables the staff to grasp the customer's emotions and conversation content in real time and immediately receive optimal product proposals and customer service methods.
[1767] "Image recognition" is a technology that detects and analyzes the characteristics of specific objects or people from image data.
[1768] "Emotion estimation" means estimating a person's emotional state (e.g., joy, anger, sadness, etc.) based on image and audio data.
[1769] "Speech recognition" is a technology that analyzes voice data and converts it into text data.
[1770] A "large-scale language model" is an artificial intelligence model trained on vast amounts of text data that is capable of understanding and generating natural language.
[1771] An "augmented reality (AR) device" is a device that displays computer graphics overlaid on images of the real world.
[1772] "Analyzing facial expressions to obtain facial emotion data" means using image recognition technology to extract emotional information from a person's facial expressions and obtain that data.
[1773] "Analyzing voice data and converting it into text" means converting voice data into text data using voice recognition technology.
[1774] "Generating appropriate product proposals by combining text data and facial emotion data" means automatically generating optimal product proposals for customers using text data and facial emotion data.
[1775] "Synthesize voice and notify staff" means synthesizing the generated product suggestions and advice into voice and conveying it to staff.
[1776] This invention provides a system that makes full use of image recognition, voice recognition, and large-scale language models to support customer service in brick-and-mortar stores. Specific embodiments for implementing this system are described below.
[1777] System configuration
[1778] The system consists of the following main components:
[1779] 1. Client devices (including smart glasses and head-mounted displays)
[1780] 2. Server (located on the cloud and responsible for main processing)
[1781] 3. User interface (smartphone app or web interface)
[1782] Program Overview
[1783] Emotion estimation using image recognition
[1784] The client device captures the customer's facial expressions in real time while serving customers in a physical store and sends the image data to a server. The server then uses image recognition technology such as Amazon Rekognition to analyze the customer's emotions (e.g., joy, anger, sadness, etc.) and obtains facial emotion data.
[1785] Emotion estimation using speech recognition
[1786] The client device records the conversation with the customer and sends the audio data to the server. The server then converts the audio data into text data using voice recognition technology such as the Google Speech-to-Text API. This text data is then used to estimate emotions based on the conversation content.
[1787] Generating optimal suggestions using large-scale language models
[1788] The server uses a large-scale language model (e.g., GPT-4) to generate optimal product suggestions and customer service methods based on text data and facial emotion data. The generated suggestions are input as prompt sentences as follows:
[1789] Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation: "What products are you interested in recently?"
[1790] Display suggestions and speech synthesis
[1791] The generated suggestions are notified to staff using voice synthesis technology (such as pyttsx3). The user (staff member) receives this information in real time through smart glasses or a head-mounted display attached to the client device and can take appropriate action.
[1792] Specific examples
[1793] Example 1: Proposing a new product
[1794] When a customer asks a staff member, "What products do you recommend?", the system converts the voice data into text data and analyzes facial expressions to infer that the customer is "interested." Based on this information, a large-scale language model generates a suggestion such as "Recommend a newly arrived product," and notifies the staff member via voice synthesis.
[1795] Example 2: Follow-up suggestions
[1796] After the staff member has finished serving the customer, the system analyzes the acquired data and generates an appropriate follow-up action (e.g., sending a thank-you message) and notifies the user interface. The user can then confirm the suggestion and take action if necessary.
[1797] As a result, this system allows staff to understand customers' emotions and conversation content in real time, enabling them to instantly receive optimal product suggestions and customer service methods.
[1798] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1799] Step 1:
[1800] The user wears the client terminal.
[1801] How it works: The customer wears smart glasses or a head-mounted display, which allows the device's camera and microphone to capture the customer's video and audio in real time.
[1802] Input: The client device worn by the user.
[1803] Output: Real-time video and audio capture.
[1804] Step 2:
[1805] The terminal captures the customer's video and audio and sends them to the server.
[1806] How it works: The device's camera captures the customer's facial expressions, and the microphone records the conversation. This data is then sent to a cloud server.
[1807] Input: Video and audio data acquired by the device.
[1808] Output: Video and audio data sent to the server.
[1809] Step 3:
[1810] The server uses image recognition technology to analyze the customer's facial expressions and obtain facial emotion data.
[1811] How it works: The server runs an image recognition algorithm such as Amazon Rekognition to estimate the customer's emotions from their facial expressions. For example, the results can be "happy" or "interested" based on the facial expression.
[1812] Input: Video data sent to the server.
[1813] Output: Facial emotion data (e.g., "HAPPY", confidence level 95.0%).
[1814] Step 4:
[1815] The server uses voice recognition technology to convert the voice data into text data.
[1816] What happens: The server runs a speech recognition algorithm, such as the Google Speech-to-Text API, to convert the conversation into text. For example, it might generate something like, "What products are you interested in these days?"
[1817] Input: The audio data sent to the server.
[1818] Output: Text data (e.g., "What products have you been interested in recently?").
[1819] Step 5:
[1820] The server combines text data and facial emotion data to input prompts into a large-scale language model to generate optimal product suggestions.
[1821] Specific operation: The server uses a large-scale language model (e.g., GPT-4) to generate a prompt sentence and then generates appropriate product suggestions based on it. An example of a prompt sentence is "Facial expression: [{'Type': 'HAPPY', 'Confidence': 95.0}], Conversation content: 'What products are you interested in recently?'"
[1822] Input: Text data and facial emotion data.
[1823] Output: Product suggestions (e.g. "Recommend new products").
[1824] Step 6:
[1825] The product suggestions generated by the server are notified to staff using voice synthesis technology.
[1826] Specific operation: The server uses a speech synthesis algorithm such as pyttsx3 to convert the generated product suggestions into voice and send it to the client terminal. The user (staff member) receives the suggestions by voice and can respond appropriately.
[1827] Input: Server-generated product suggestions.
[1828] Output: A voice-synthesized product suggestion notification.
[1829] Step 7:
[1830] Based on the user-generated proposals, appropriate products are proposed to the customer.
[1831] Specific operation: The user (staff member) uses the product proposal notified by voice from the client terminal as a reference to explain and propose new products to the customer.
[1832] Input: Speech-synthesized product suggestions.
[1833] Output: Product recommendations to customers.
[1834] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1835] This invention is a system that analyzes the emotions of users and their partners in real time during dates and business meetings, and suggests optimal topics and behaviors. This system achieves highly accurate emotion analysis by combining image recognition, speech recognition technology, large-scale language models, and an emotion engine.
[1836] Overall system configuration
[1837] The system consists of the following main components:
[1838] 1. Client device (augmented reality (AR) device worn by the user)
[1839] 2. Server (responsible for major processing on the cloud)
[1840] 3. Emotion engine (analyzing user and other party emotions in real time)
[1841] 4. User interface (smartphone app or web interface)
[1842] Initial setup and preparation phase
[1843] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate optimal initial topics and suggestions for the other person, and propose date plans and sales strategies. The user selects a proposed plan and sends it to the server, completing the reservation arrangements.
[1844] Real-time support during business hours
[1845] When a user wears AR glasses or AR contact lenses, the device captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the other person's emotions. At the same time, an emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[1846] The server integrates the emotional information of both the user and the other party, analyzes the current dating or sales context using a large-scale language model, and generates optimal advice. The generated advice is displayed in the user's field of vision, and the user can use it as a reference to continue the conversation.
[1847] Furthermore, the emotion engine learns from the user's past emotional data, analyzes their emotional tendencies, and provides advice in response to predicted emotional changes. This makes it easier for the user to understand their own emotional state and continue the conversation calmly.
[1848] Follow-up
[1849] In the follow-up stage, the user notifies the app that the date or business has ended. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending a message), and notifies the user. The notified message can be sent after the user confirms it.
[1850] Specific examples
[1851] In the case of a date
[1852] 1. The day before the date
[1853] The user enters the other person's basic information into a smartphone app.
[1854] The server suggests plans such as "lunch at Cafe A" or "watching Movie B."
[1855] The user selects "Watch Movie B" and the server reserves the tickets.
[1856] 2. The day of the date
[1857] The user puts on the AR glasses.
[1858] During the date, the AR glasses transmit video and audio to a server.
[1859] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[1860] A large-scale language model generates advice like, "Talk about a movie you recently saw."
[1861] The advice is displayed in the user's field of view, and the user provides the topic.
[1862] 3. After the date
[1863] The server generates a message saying, "Thank you for today. It was fun!"
[1864] The message is notified to the smartphone app, and the user confirms it before sending it.
[1865] In the case of sales
[1866] 1. Before the sales visit
[1867] The user enters the other company's information and the name of the person in charge into the smartphone app.
[1868] The server prepares initial proposals and topics based on past transaction history and related information.
[1869] 2. During a sales visit
[1870] The user puts on the AR contact lenses.
[1871] During business hours, the AR contact lenses capture video and audio and send them to a server.
[1872] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[1873] The large-scale language model generated the advice "Introduce and demonstrate the new product."
[1874] The advice is displayed in the user's field of view and the user demonstrates it.
[1875] 3. After the visit
[1876] The server generates a follow-up email.
[1877] A follow-up email is sent to the smartphone app and sent after confirmation by the user.
[1878] In this way, the present invention is a system that improves the quality of communication by analyzing the emotions of the user and the other party in real time and suggesting optimal topics and actions. The introduction of an emotion engine makes it easier to understand the user's own emotions, enabling more effective support.
[1879] The processing flow will be explained below.
[1880] Initial setup and preparation phase
[1881] Step 1:
[1882] The user launches the smartphone app, enters their login information, and then enters the other person's basic information (such as name, age, hobbies, and previous conversations) into the app.
[1883] Step 2:
[1884] The server receives the basic information of the other party sent by the user, searches the database of past conversation history based on the received information, and retrieves matching data.
[1885] Step 3:
[1886] The server uses a large-scale language model to generate optimal initial topics and suggestions for the other party, and notifies the user of the generated suggestions via their smartphone app.
[1887] Step 4:
[1888] The user checks the proposed date plans and sales strategies and selects one. The selected information is sent to the server.
[1889] Step 5:
[1890] The server makes the necessary reservations (for restaurants, cinemas, etc.) based on the user's selection, and notifies the user's smartphone app that the reservation has been completed.
[1891] Real-time support during business hours
[1892] Step 1:
[1893] The user puts on the AR glasses or AR contact lenses and checks the device pairing in the smartphone app.
[1894] Step 2:
[1895] The device uses its sensors to capture video and audio in real time, and the captured data is sent to a server.
[1896] Step 3:
[1897] The server analyzes the video data it receives using an image recognition algorithm to infer emotions from facial expressions, and it also analyzes audio data using a voice recognition algorithm to infer emotions from the tone and patterns of voice.
[1898] Step 4:
[1899] The server creates the other person's emotional information based on the emotion estimation results. At the same time, the emotion engine analyzes the user's emotions in real time and monitors their emotional state.
[1900] Step 5:
[1901] The server integrates the emotional information of both the user and the other party and generates optimal advice using a large-scale language model. The advice is then displayed in the user's field of view.
[1902] Step 6:
[1903] The device displays the advice received from the server in the user's field of view, and the user continues the dialogue while referring to the displayed advice.
[1904] Follow-up
[1905] Step 1:
[1906] The user enters into the app that a date or business trip has ended, and the device sends all data up to the end to the server.
[1907] Step 2:
[1908] The server analyzes the data at the end and generates a recommended action to be taken next, which is then sent to the user's smartphone app.
[1909] Step 3:
[1910] The user reviews the follow-up message suggested by the app, edits it if necessary, and sends it via LINE or email.
[1911] Example 2
[1912] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1913] In conventional dating and sales situations, it has been difficult for users to accurately read the other person's emotions and decide on appropriate topics and actions. Furthermore, it has been difficult for users to manage their own emotions while having an effective conversation. This has led to a decline in the quality of communication and a decrease in the success rate of dates and sales.
[1914] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through voice recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and an emotion engine for analyzing the user's emotions in real time. This configuration enables the user to grasp the emotions of others and themselves in real time and take optimal actions and responses.
[1915] "Image recognition" is the technology that extracts information from digital images and videos and identifies specific patterns and features.
[1916] "Speech recognition" is a technology that analyzes voice data and converts spoken words into text data.
[1917] A "large-scale language model" is an artificial intelligence model trained on large amounts of text data, and is a model that excels at understanding and generating natural language.
[1918] An "augmented reality (AR) device" is a device that overlays digital information onto a user's real-world field of vision. Examples include AR glasses and AR contact lenses.
[1919] The "emotion engine" is a system that analyzes the emotional state of the user and the other person in real time and provides appropriate feedback and advice based on the results.
[1920] A "date plan" refers to a proposal or action plan for systematically preparing the progress of a date.
[1921] A "sales strategy" is a plan for effective contact methods and presentations in sales activities.
[1922] "Real-time data" is data that represents ongoing events or conditions and should be processed immediately.
[1923] "Follow-up actions" are subsequent actions that should be taken after a date or business meeting, with the purpose of maintaining and strengthening the relationship with the other person.
[1924] "Initial settings" refers to the setup work required before starting to use the system, in which basic information about the user and the other party is entered and preparations required for system operation are made.
[1925] The system based on this invention analyzes the emotions of the user and the other party in real time during dating or business, and suggests optimal topics and behaviors. This system includes the following main components: image recognition, speech recognition, a large-scale language model, an augmented reality (AR) device, and an emotion engine.
[1926] Hardware and software used
[1927] Hardware: Augmented reality (AR) glasses, augmented reality (AR) contact lenses, smartphones
[1928] Software: Image recognition software (e.g., Google Cloud Vision), speech recognition software (e.g., Google Speech-to-Text), large-scale language models (e.g., OpenAI GPT-3), sentiment analysis engines (e.g., Affectiva)
[1929] Initial Setup and Data Entry
[1930] The user launches the smartphone app and enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives the basic information sent by the user, coordinates it with past conversation history data to perform initial setup, and generates optimal initial topics and suggestions (e.g., date plans or sales strategies). The user selects a suggested plan on the smartphone app, and the server arranges reservations and other preparations based on the selected plan.
[1931] Real-time data acquisition and analysis
[1932] The user wears an AR device (e.g., AR glasses or AR contact lenses), and the device captures video and audio in real time. The captured data is sent to a server, which analyzes it using image and voice recognition software. The analysis results are processed by an emotion engine, which analyzes the emotional state of the other person and the user themselves.
[1933] Generating and displaying advice
[1934] Based on the analysis results, the server uses a large-scale language model to generate optimal topics and behaviors. The generated advice is sent to the user's AR device and displayed as an overlay in the user's field of view. The user can refer to the displayed advice to proceed with the conversation.
[1935] Follow-up
[1936] After a date or business meeting ends, the user notifies the end of the event via a smartphone app. The server analyzes all data up to the end of the event and generates the next follow-up action (such as sending a message). The generated message is notified to the user, who can then send it after confirming it.
[1937] Specific examples
[1938] Specific examples of dating
[1939] 1. The day before the date
[1940] The user enters the other person's basic information into a smartphone app.
[1941] The server will suggest plans such as "lunch at a cafe" or "watching a movie."
[1942] The user selects "Watch a Movie" and the server reserves the tickets.
[1943] 2. The day of the date
[1944] The user puts on the AR glasses.
[1945] During the date, the AR glasses transmit video and audio to a server.
[1946] The server analyzes the other person's facial expressions and tone of voice, and the emotion engine monitors the user's emotional state.
[1947] A large-scale language model generates advice like, "Tell me about a movie you recently saw."
[1948] The advice is displayed in the user's field of view, and the user provides the topic.
[1949] 3. After the date
[1950] The server generates a message saying "Thanks for today, see you again soon!"
[1951] The message is notified to the smartphone app, and the user confirms it before sending it.
[1952] Specific examples of sales
[1953] 1. Before the sales visit
[1954] The user enters the other company's information and the name of the person in charge into the smartphone app.
[1955] The server references transaction history and related information to prepare initial proposals and topics.
[1956] 2. During a sales visit
[1957] The user puts on the AR contact lenses.
[1958] During business hours, the AR contact lenses capture video and audio and send it to a server.
[1959] The server analyzes the facial expressions and tone of voice of the person in charge, and the emotion engine monitors the user's emotional state.
[1960] A large-scale language model generates the advice "Introduce and demonstrate the new product."
[1961] The advice is displayed in the user's field of view and the user demonstrates it.
[1962] 3. After business hours
[1963] The server generates a follow-up email to notify the user.
[1964] The user checks the email and then sends it.
[1965] Prompt Sentence Examples
[1966] "If the other person seems to be having fun, what topic should I bring up next?"
[1967] "What should I do when my partner seems a little cranky?"
[1968] "Suggest topics based on topics that your partner has shown interest in on previous dates."
[1969] In this way, the system based on this invention analyzes the emotions of the user and the other party in real time and suggests optimal topics and actions, thereby improving the quality of communication. The introduction of an emotion engine makes it easier for users to understand their own emotions, enabling more effective support.
[1970] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1971] Step 1:
[1972] The user launches the smartphone app. The user enters the other person's basic information (such as name, age, hobbies, and previous conversations). The server receives this basic information and performs initial setup in conjunction with past conversation history data. Specifically, the server stores the entered name and hobbies of the other person in a database and compares them with past conversation history data. As an output, the most suitable initial topic and suggestions for the other person are generated based on the comparison results.
[1973] Step 2:
[1974] The user selects a proposed plan (e.g., a date plan or a sales strategy) on a smartphone app. The server receives the selected plan and arranges and prepares the necessary reservations based on that plan. At this time, the server connects to the reservation system and secures the necessary resources (e.g., movie tickets, restaurant reservations). As an output, the confirmed date plan or sales strategy is generated and notified to the user.
[1975] Step 3:
[1976] A user puts on an augmented reality (AR) device (e.g., AR glasses or AR contact lenses). The device captures video and audio in real time. This data is collected through the device's camera and microphone. The input real-time data is sent to a server, which prepares it for data analysis. Specifically, video data is sent to the server in streaming format, and audio data is transferred in real time. Data begins to be sent to the server as output.
[1977] Step 4:
[1978] The server analyzes the video data using image recognition software (e.g., Google Cloud Vision) and the audio data using voice recognition software (e.g., Google Speech-to-Text). In this analysis, the other person's facial expressions and tone of voice are extracted as important elements. The input is video and audio data transmitted in real time, and the output is data indicating the emotional state. Specifically, the server extracts facial expression features from the video data and extracts tone of voice and word features from the audio data.
[1979] Step 5:
[1980] The emotion engine analyzes the emotional state of the user and the other party in real time based on the analysis results. The input is the analysis results of image recognition and voice recognition, and the output is emotion analysis data. The emotion engine integrates this data and grasps the psychological state of the user and the other party in real time. Specifically, the emotion engine also monitors changes in the user's heart rate and breathing.
[1981] Step 6:
[1982] The server inputs the sentiment analysis data into a large-scale language model (e.g., OpenAI GPT-3) to generate optimal topics and behaviors. The input is the sentiment analysis data and provided context information, and the output is specific advice. By sending prompts to the large-scale language model, responses are generated to questions such as, "If the other person seems to be having fun, what topic should I bring up next?"
[1983] Step 7:
[1984] The server compiles the generated advice and sends it to the user's AR device. The input is advice generated by a large-scale language model, and the output is information overlaid on the user's field of view. Specifically, the server sends appropriate overlay information to the AR device, and the advice is displayed in the user's field of view.
[1985] Step 8:
[1986] The user proceeds with the dialogue while referring to the displayed advice. The input is the advice displayed on the AR device, and the output is the user's actions. The user interacts with the other party based on the advice, and the results are sent to the server as data acquired in real time. Specifically, the dialogue content and the other party's reactions are continuously acquired in real time.
[1987] Step 9:
[1988] After a date or business meeting ends, the user notifies the smartphone app that the event is over. The server analyzes all data up to the end and generates the follow-up action to be taken next. The input is all data up to the end, and the output is the follow-up action (e.g., sending a message). In concrete terms, the server comprehensively analyzes all the data obtained and determines the next step.
[1989] Step 10:
[1990] The server notifies the user of a follow-up message generated by the server. The user confirms and sends the message. The input is the message generated by the server, and the output is the message confirmed and sent by the user. Specifically, the user confirms the message in the smartphone app, edits or corrects it as necessary, and then taps the send button.
[1991] (Application example 2)
[1992] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1993] In conventional marketing activities, it has been difficult to instantly grasp the emotional reactions of clients and customers and propose optimal advertising strategies in real time. This has led to problems such as not being able to maximize the effectiveness of advertising and insufficient communication with clients. In particular, it has been difficult to provide effective feedback in situations where it is necessary to respond quickly to client reactions during a conversation.
[1994] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1995] In this invention, the server includes a means for estimating emotions through image recognition, a means for estimating emotions through speech recognition, a means for generating optimal topics and behaviors using a large-scale language model, a means for displaying information on an augmented reality (AR) device worn by the user, and a means for analyzing the emotional responses of conversation partners in real time during marketing activities and generating optimal advertising strategies and proposals. This makes it possible to analyze the emotions of clients or customers in real time and quickly generate optimal advertising strategies and proposals accordingly. This system is effective for marketers to receive instant feedback during conversations and promote effective communication with clients.
[1996] "Image recognition" is a technology that extracts and analyzes specific information and features from image data acquired using devices such as cameras.
[1997] "Means for estimating emotions" refers to technology that uses image recognition and voice recognition to analyze a person's facial expressions and tone of voice, and then estimates emotions based on the results.
[1998] "Speech recognition" is a technology that analyzes voice data acquired through devices such as microphones and extracts characteristics of words and voices.
[1999] A "large-scale language model" is a machine learning model trained using massive amounts of text data, and has the ability to understand and generate natural language.
[2000] "Means for generating optimal topics and behaviors" refers to a technology that uses a large-scale language model to generate suggestions for optimal topics and actions according to the situation.
[2001] A "user-worn augmented reality (AR) device" is a device worn by a user that can overlay digital information on the user's field of vision.
[2002] "Marketing activities" refers to various sales activities aimed at promoting products and services.
[2003] "Means for analyzing the emotional responses of a conversation partner in real time" refers to technology that instantly analyzes the emotions and responses of the other person during a conversation and provides that information.
[2004] An "advertising strategy" refers to a plan or method for effectively promoting a particular product or service.
[2005] The "means for generating suggestions" is a technology that suggests optimal actions and conversation content to users based on analyzed data.
[2006] The present invention provides a system for analyzing the emotions of a conversation partner in real time during marketing activities and for optimizing advertising strategies and proposals. The system is realized using image recognition, speech recognition, a large-scale language model, a wearable augmented reality (AR) device, and an emotion analysis engine. Specific embodiments for implementing the present invention are described below.
[2007] Overall system configuration
[2008] The system consists of the following main components:
[2009] 1. Client device (augmented reality (AR) device worn by the user)
[2010] 2. Server (responsible for major processing on the cloud)
[2011] 3. Emotion engine (analyzing the emotions of the person interacting with you and the user in real time)
[2012] 4. User interface (smartphone app or web interface)
[2013] Initial setup and preparation phase
[2014] First, the user launches the smartphone app and registers the other party's basic information (such as name, age, hobbies, and response to previous advertisements) in natural language. The server receives the information sent by the user and performs initial setup in conjunction with past conversation history data. This allows the server to generate an optimal initial advertising strategy and proposal for the other party and propose a marketing plan. The user selects the proposed plan and sends it to the server, completing preparations for marketing activities.
[2015] Real-time support during marketing activities
[2016] The user wears an augmented reality (AR) device, which captures video and audio in real time. The captured data is sent to a server, which uses image and voice recognition to estimate the emotion of the person they are interacting with. At the same time, an emotion engine analyzes and monitors the user's emotional state in real time.
[2017] The server integrates the emotional information of the user and the conversation partner, analyzes the context of current marketing activities using a large-scale language model, and generates optimal advertising strategies and proposals. The generated proposals are displayed in the user's field of view, and the user continues the dialogue using them as reference. Furthermore, the emotion engine learns the user's past emotional data, analyzes their emotional tendencies, and makes proposals that correspond to predicted emotional changes.
[2018] Follow-up
[2019] After the marketing activity is completed, the user notifies the smartphone app. The server analyzes all data up to the end, generates the next follow-up action to be taken (e.g., sending an email), and notifies the user. The notified message can be sent after the user confirms it.
[2020] Hardware and software used
[2021] Hardware:
[2022] Smartphone
[2023] camera
[2024] Augmented Reality (AR) Devices
[2025] software:
[2026] OpenCV: Image Recognition Library
[2027] dlib: face detection library
[2028] DeepFace: A library for emotion analysis
[2029] OpenAI API: Large-scale language models
[2030] Specific examples
[2031] During a meeting with an important client, a marketing manager at an advertising agency uses a smartphone camera to analyze the client's facial expressions. The manager and his colleagues then review the emotion analysis results to devise strategies and make optimal proposals in real time.
[2032] Prompt Sentence Examples
[2033] My client is expressing "delight" emotions. What advertising strategy should I suggest next?
[2034] My client is expressing "anger" emotions. How do I fix this next?
[2035] My client is expressing feelings of sadness. How should I follow up next?
[2036] As described above, the present invention provides a system that combines sentiment analysis and large-scale language models to provide optimal advertising strategies and proposals in real time for marketing activities.
[2037] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2038] Step 1:
[2039] The server receives basic information about the other party (such as name, age, hobbies, and response to previous advertisements) from the user's device (smartphone app). This allows the necessary information to be accumulated on the server. Natural language data entered on the user's device is sent to the server, where it is analyzed and saved.
[2040] Step 2:
[2041] The server references the received contact's basic information and past conversation history data to perform initial setup. It also generates and proposes an appropriate initial advertising strategy and marketing plan. The plan is generated using a large-scale language model, using the contact's basic information and past conversation history data as input.
[2042] Step 3:
[2043] The user selects a proposed plan and sends it to the server. The server determines a marketing activity plan based on the selected plan and notifies the user's device. The selected plan and related data are processed and stored on the server.
[2044] Step 4:
[2045] A user wears an augmented reality (AR) device and starts interacting with the client. The AR device captures video and audio in real time and sends them to the server. The video and audio data are sent to the server as input.
[2046] Step 5:
[2047] The server analyzes the received video data using OpenCV and dlib to perform facial recognition and emotion estimation (image recognition). At the same time, it analyzes the received audio data using DeepFace to estimate emotions (speech recognition). The results of the emotion analysis are generated and saved on the server.
[2048] Step 6:
[2049] The emotion analysis engine monitors the emotional state of the user and the person they are interacting with based on the analyzed emotion data. The monitoring data is created on the server.
[2050] Step 7:
[2051] It uses a large-scale language model to analyze the context of current marketing activities and generate optimal advertising strategies and proposals. Using prompt sentences, it generates specific advertising strategies and proposals and displays them in the user's field of view. The generated advertising strategies and proposals are displayed on the user's AR device.
[2052] Step 8:
[2053] The user visually confirms the suggestions through the AR device and applies them during the dialogue to communicate with the client. The user advances the conversation based on the generated strategies and suggestions.
[2054] Step 9:
[2055] After completing a marketing activity, the user notifies the server using the smartphone app.
[2056] Step 10:
[2057] The server comprehensively analyzes all log data and sentiment analysis data and generates the next follow-up action to be taken. The generated action and the content of the follow-up are notified to the user's device.
[2058] Step 11:
[2059] The user confirms the follow-up action (e.g., sending an email) and sends it as is if necessary. The follow-up action is executed from the user's device.
[2060] Through the above steps, the system of the present invention can analyze the emotions of the person in conversation in real time during marketing activities and make optimal advertising strategies and proposals.
[2061] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2062] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2063] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2064] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2065] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2066] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2067] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2068] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2069] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2070] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2071] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2072] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2073] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2074] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2075] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2076] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2077] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2078] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2079] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2080] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2081] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2082] The following is further disclosed regarding the above embodiment.
[2083] (Claim 1)
[2084] A means for estimating emotions through image recognition;
[2085] means for estimating emotion through speech recognition;
[2086] A means of generating optimal topics and behaviors using a large-scale language model;
[2087] means for displaying information on an augmented reality (AR) device worn by a user;
[2088] A system including:
[2089] (Claim 2)
[2090] A means for a user to register basic information of a contact by natural language input;
[2091] A means for initial setup in conjunction with past conversation history data;
[2092] A way to propose date plans and sales strategies,
[2093] The system of claim 1 further comprising:
[2094] (Claim 3)
[2095] A way for users to receive follow-up actions after a date or business meeting;
[2096] A way to generate LINE and email messages,
[2097] The system of claim 1 further comprising:
[2098] "Example 1"
[2099] (Claim 1)
[2100] A means for estimating emotions through image recognition;
[2101] means for estimating emotion through speech recognition;
[2102] A means of generating optimal topics and behaviors using a large-scale language model;
[2103] means for displaying information on an augmented reality (AR) device worn by a user;
[2104] A means for capturing video and audio in real time and transmitting them to a server;
[2105] A means to analyze all data up to the end and generate the next action;
[2106] A system including:
[2107] (Claim 2)
[2108] A means for a user to register basic information of a contact by natural language input;
[2109] A means for initial setup in conjunction with past conversation history data;
[2110] A way to propose date plans and sales strategies,
[2111] means for the user to select a proposed plan and transmit it to the server;
[2112] A means for the server to arrange reservations required for the selected plan;
[2113] The system of claim 1 further comprising:
[2114] (Claim 3)
[2115] A way for users to receive follow-up actions after a date or business meeting;
[2116] A means for generating and notifying a message text;
[2117] The system of claim 1 further comprising:
[2118] "Application Example 1"
[2119] (Claim 1)
[2120] A means for estimating emotions through image recognition;
[2121] means for estimating emotion through speech recognition;
[2122] A means of generating optimal topics and behaviors using a large-scale language model;
[2123] means for displaying information on an augmented reality (AR) device worn by a user;
[2124] A means of analyzing customer facial expre...
Claims
1. A means for estimating emotions through image recognition; means for estimating emotion through speech recognition; A means of generating optimal topics and behaviors using a large-scale language model; means for displaying information on an augmented reality device worn by a user; A system including:
2. A means for a user to register basic information of a contact by natural language input; A means for initial setup in conjunction with past conversation history data; A way to propose date plans and sales strategies, The system of claim 1 further comprising:
3. A way for users to receive follow-up actions after a date or business meeting; A way to generate LINE and email messages, The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A