System

The communication support system addresses the challenge of reduced face-to-face communication by using real-time voice and image analysis to suggest honorifics and answers, improving communication effectiveness in various settings.

JP2026023906APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126227
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Individuals, particularly those in their 40s to 60s, face challenges in using appropriate honorific language and vocabulary due to reduced face-to-face communication opportunities and increased remote work, leading to difficulties in business and daily interactions.

Method used

A communication support system that utilizes real-time analysis of user voice and image data through a generative AI model to provide appropriate honorific expressions, word suggestions, location descriptions, and answers to questions, using devices like smart glasses or smartwatches to facilitate effective communication.

Benefits of technology

Enables users to maintain appropriate communication skills and reduce difficulties in business and daily life by providing immediate and accurate suggestions and information, enhancing communication quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023906000001_ABST
    Figure 2026023906000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A communication support system comprising: means for acquiring voice data of a user in real time; means for transmitting the voice data to a server; means for analyzing the voice data by the server and generating a proposal of an appropriate honorific or word; and means for feeding back the generated proposal to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The purpose of this invention is to support users who have difficulty choosing words and using honorific language due to the decrease in opportunities for face-to-face communication caused by the aging population and the increase in remote work. In particular, there is a need to provide a means to provide quick and appropriate advice for forgetfulness and language problems, which are common among users in their 40s to 60s. [Means for solving the problem]

[0005] The present invention provides a communication support system including: means for acquiring a user's voice data in real time; means for transmitting the voice data to a server; means for the server to analyze the voice data and generate appropriate honorific expressions and word suggestions; and means for providing the generated suggestions to the user. The system further includes means for acquiring a user's image data; means for transmitting the image data to a server; means for the server to analyze the image data and generate location descriptions and object names; and means for providing the generated information to the user. The system also includes means for acquiring question data from the user; means for transmitting the question data to a server; means for the server to analyze the question data and generate answers; and means for providing the generated answers to the user. In this way, users can maintain appropriate communication skills and reduce difficulties in business and daily life.

[0006] "User" refers to a person who uses the system and has the role of providing specific audio and image data.

[0007] "Voice data" refers to information that is a digital recording of a user's spoken words or voice.

[0008] "Image data" refers to a digital record of visual information seen or captured by a user.

[0009] "Real-time" refers to a situation where data acquisition and processing occurs immediately without delay.

[0010] "Server" refers to a central processing unit for analyzing data and generating processed results.

[0011] "Honorific language" refers to the Japanese language used to express politeness when speaking to someone.

[0012] "Word suggestions" refers to presenting appropriate words and expressions based on what the user is saying.

[0013] "Feedback" refers to information or advice provided to the user as a result of analysis or evaluation.

[0014] "Communication support system" refers to an integrated platform that supports user communication through the analysis of voice and image data. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention is a communication support system that uses AI as its underlying technology, and analyzes the user's voice and image data in real time to suggest appropriate honorific expressions and words, as well as provide answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[0037] System configuration

[0038] User

[0039] This is the entity that uses the system to receive communication support. They use devices such as smart glasses or smart watches.

[0040] Terminal

[0041] A device that captures audio and image data in real time. This includes smart glasses and smart watches. The device transmits the captured data to a server.

[0042] server

[0043] It is a central processing unit that analyzes the received voice and image data. The server uses a generative AI model to analyze the data, generate appropriate suggestions and answers, and send them back to the device.

[0044] Program processing

[0045] Acquisition and analysis of audio data

[0046] 1. The user puts on the smart glasses or smartwatch and starts a conversation.

[0047] 2. The device captures the user's voice data in real time and sends the data to the server.

[0048] 3. The server analyzes the received voice data using a generative AI model. For example, if a user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates the suggestion "It's an honor to meet you."

[0049] 4. The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[0050] Image data acquisition and analysis

[0051] 1. When a user wants to see a particular place or object, they capture an image of it through the smart glasses.

[0052] 2. The device sends the captured image data to the server.

[0053] 3. The server analyzes the received image data and generates a description of the location or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on the right."

[0054] 4. The server sends the generated information to the terminal, which provides it to the user.

[0055] Real-time question answering

[0056] 1. A user asks a specific question via voice or text, for example, "What does this word mean?"

[0057] 2. The device sends the question data to the server.

[0058] 3. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'."

[0059] 4. The server generates a response and sends it to the device, which provides it to the user via voice or text.

[0060] Specific examples

[0061] Example 1: Use during a business meeting

[0062] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[0063] Example 2: Using navigation

[0064] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[0065] As described above, the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. This system is usable by a wide range of users regardless of age or level of understanding of technology, making it useful to many people.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The user puts on the smart glasses or smartwatch and turns on the device, which starts up the device and connects it to the network.

[0069] Step 2:

[0070] The device activates its audio and image sensors and prepares to capture data in real time.

[0071] Step 3:

[0072] When a user starts a conversation and is unsure about the meaning of a word or how to use honorific language, they speak out verbally.

[0073] Step 4:

[0074] The terminal acquires the user's voice data in real time.

[0075] Step 5:

[0076] The terminal transmits the acquired voice data to the server.

[0077] Step 6:

[0078] The server analyzes the received voice data using a generative AI model. For example, if a user says, "Oh, sorry. What was it?", the server analyzes the voice data and generates an appropriate polite expression such as, "Thank you for your hard work. What are you looking for?"

[0079] Step 7:

[0080] The server sends the analysis results to the device.

[0081] Step 8:

[0082] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[0083] Step 9:

[0084] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[0085] Step 10:

[0086] The terminal transmits the captured image data to the server.

[0087] Step 11:

[0088] The server analyzes the received image data and generates information such as "The conference room is further ahead on the right."

[0089] Step 12:

[0090] The server transmits the generated information to the terminal.

[0091] Step 13:

[0092] The device provides navigation information to the user visually (on the smart glasses display).

[0093] Step 14:

[0094] The user speaks or texts a specific question, for example, "What does this word mean?"

[0095] Step 15:

[0096] The terminal transmits the question data to the server.

[0097] Step 16:

[0098] The server analyzes the question data and generates an answer such as "The meaning of this word is 'Tango'."

[0099] Step 17:

[0100] The server sends the generated response to the terminal.

[0101] Step 18:

[0102] The device provides the user with a response by voice or text.

[0103] Step 19:

[0104] The user selects a subscription plan using a dedicated application.

[0105] Step 20:

[0106] The terminal transmits the selected plan information to the server.

[0107] Step 21:

[0108] The server enables certain features based on the plan selected by the user.

[0109] Step 22:

[0110] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[0111] This allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life.

[0112] Example 1

[0113] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0114] In modern society, communication is extremely important in business and everyday life. However, many people have difficulty using appropriate honorifics and vocabulary, and recognizing places and objects, as well as answering questions instantly, are now essential needs. Conventional communication support systems lack the ability to efficiently solve these problems in real time. In particular, more advanced technology is needed to provide appropriate content suggestions and rapid analytical feedback in a variety of situations.

[0115] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0116] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to a central processing unit, means for analyzing the voice data by the central processing unit and generating suggestions for appropriate honorific expressions and words, means for feeding back the generated suggestions to the user, and a generative AI model for analyzing the voice data. As a result, the user can receive suggestions for appropriate honorific expressions and words in real time, improving the quality of communication and reducing difficulties in business and daily life.

[0117] A "user" is an entity that uses the system to receive communication support.

[0118] "Voice data" refers to data in which the user's voice is recorded in digital form.

[0119] "Real-time" means that processing occurs immediately and on the spot with minimal delay.

[0120] "Means of acquisition" refers to devices and functions for collecting audio data and image data.

[0121] A "central processing unit" is a server or computer system that performs major processing such as data analysis.

[0122] "Means of transmission" refers to the communications means or technology used to transfer acquired data to a specific destination.

[0123] "Means for analyzing" refers to software or algorithms used to analyze received data.

[0124] "Means for generating suggestions" refers to a function that generates appropriate information and advice to be provided to the user based on the analyzed data.

[0125] "Feedback means" refers to an output device or method for notifying the user of the generated suggestions.

[0126] "Image data" is data that records a user's visual information in digital form.

[0127] "Location description" refers to content that provides information or directions to a specific location.

[0128] "Name of object" refers to the name of the object contained in the image data.

[0129] "Question data" refers to data that indicates the content of a question that a user asks the system.

[0130] "Means for generating answers" refers to algorithms or software for generating optimal answers based on question data.

[0131] A "generative AI model" is a model or system that uses artificial intelligence to analyze data and generate suggestions or answers.

[0132] This invention is a communication support system that uses a generative AI model as its underlying technology. It analyzes the user's voice and image data in real time, and suggests appropriate honorific expressions and words, as well as providing answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[0133] Acquisition and analysis of audio data

[0134] A user puts on smart glasses or a smartwatch and starts a conversation. The device captures the user's voice data in real time and sends it to a server. The server analyzes the received voice data using a generative AI model (e.g., OpenAI's GPT-4). For example, if the user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates a suggestion: "It's an honor to meet you." The server sends the suggestion to the device, which then provides visual or auditory feedback to the user.

[0135] Image data acquisition and analysis

[0136] When a user wants to check a specific place or object, they capture an image of it through the smart glasses. The device sends the captured image data to a server. The server analyzes the received image data and generates a description of the place or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on your right." The server sends the generated information to the device, which then provides it to the user.

[0137] Real-time question answering

[0138] The user asks a specific question by voice or text. For example, "What does this word mean?", the device sends the question data to the server. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[0139] Specific examples

[0140] Example 1: Use during a business meeting

[0141] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[0142] Example 2: Using navigation

[0143] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[0144] Prompt Sentence Examples

[0145] Use the following prompt:

[0146] User Input: "Nice to meet you."

[0147] AI Model: "It's a pleasure to meet you."

[0148] User Input: "What does this word mean?"

[0149] AI Model: "This word means 'Tango'"

[0150] In this way, the system of the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. Furthermore, by using a generative AI model, it is possible to respond quickly and accurately to user needs, making the system useful for many people, regardless of age or level of technological literacy.

[0151] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0152] Acquisition and analysis of audio data

[0153] Step 1: The user puts on the smart glasses or smartwatch and starts a conversation.

[0154] Input: User's voice

[0155] Action: The user speaks.

[0156] Output: Audio signal

[0157] Step 2: The terminal acquires the user's voice data in real time.

[0158] Input: Audio signal

[0159] How it works: The device's microphone picks up audio signals and converts them into digital audio data.

[0160] Output: Digital audio data

[0161] Step 3: The device sends the acquired voice data to the server.

[0162] Input: Digital audio data

[0163] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[0164] Output: Compressed and encrypted audio data

[0165] Step 4: The server analyzes the received voice data using the generative AI model.

[0166] Input: Compressed and encrypted audio data

[0167] How it works: The server decodes the audio data and uses a generative AI model (e.g., GPT-4) to convert the audio to text and perform analysis.

[0168] Output: Analysis results (suggestions for appropriate honorifics and words)

[0169] Step 5: The server sends the generated proposal to the device.

[0170] Input: Analysis results (suggestions for appropriate honorifics and words)

[0171] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[0172] Output: Compressed and encrypted proposal data

[0173] Step 6: The device provides visual or auditory feedback to the user.

[0174] Input: Compressed and encrypted proposal data

[0175] How it works: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[0176] Output: User feedback (visual / auditory information)

[0177] Image data acquisition and analysis

[0178] Step 1: When a user wants to see a specific place or object, they capture an image of it through their smart glasses.

[0179] Input: Visual information

[0180] Action: The user operates the smartglasses camera to capture an image.

[0181] Output: Image data

[0182] Step 2: The device sends the captured image data to the server.

[0183] Input: Image data

[0184] Operation: The terminal's communication module compresses and encrypts the image data and sends it to the server.

[0185] Output: Compressed and encrypted image data

[0186] Step 3: The server analyzes the received image data.

[0187] Input: Compressed and encrypted image data

[0188] How it works: The server decodes the data and uses image recognition software to analyze the image (e.g., recognize location descriptions and object names).

[0189] Output: Analysis results (location descriptions and names of things)

[0190] Step 4: The server sends the generated information to the terminal.

[0191] Input: Analysis results (location description and object names)

[0192] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[0193] Output: Compressed and encrypted information data

[0194] Step 5: The terminal provides the user with information visually or audibly.

[0195] Input: Compressed and encrypted information data

[0196] How it works: The device decodes the data and displays it on the screen or speaks the information over the speaker.

[0197] Output: Provide information to the user (visual / auditory information)

[0198] Real-time question answering

[0199] Step 1: The user asks a specific question via voice or text.

[0200] Input: Voice or text question

[0201] How it works: If the user asks a question by voice, the device's microphone picks up the sound, and if the user asks a question by text, the device's text input function is used.

[0202] Output: Digital audio data or text data

[0203] Step 2: The device sends the query data to the server.

[0204] Input: Digital audio or text data

[0205] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[0206] Output: Compressed and encrypted query data

[0207] Step 3: The server analyzes the question data and generates the best answer.

[0208] Input: Compressed and encrypted query data

[0209] How it works: The server decodes the data, uses a generative AI model to analyze the intent of the question, and generates an answer.

[0210] Output: Analysis results (best answer)

[0211] Step 4: The server sends the generated response to the terminal.

[0212] Input: Analysis results (best answer)

[0213] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[0214] Output: Compressed and encrypted answer data

[0215] Step 5: The device provides the user with a voice or text message.

[0216] Input: Compressed and encrypted answer data

[0217] What it does: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[0218] Output: Provides answers to users (visual / auditory information)

[0219] As described above, the system achieves efficient communication support for users through specific operations and data input / output at each step.

[0220] (Application example 1)

[0221] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0222] Conventional communication support systems allow users to receive suggestions for appropriate honorific expressions and words, but they have difficulty answering customer questions in real time or providing in-store navigation information when dealing with customers in brick-and-mortar stores. They also lack the ability to provide personalized services based on customer facial recognition or past purchase history. The present invention provides a new communication support system that streamlines customer service in brick-and-mortar stores and improves customer satisfaction.

[0223] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0224] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, means for analyzing question data acquired from the user via the smart wearable device and generating an optimal answer using a generative AI model, means for visually or audibly feeding back the generated answer, and means for displaying in-store navigation information based on the generated information, thereby enabling immediate responses to customer questions in physical stores, in-store navigation, and personalized services.

[0225] "Voice data" is information about sounds generated by the user's vocalizations.

[0226] A "server" is a central processing unit that processes and provides various data via a network.

[0227] "Analysis" is the act of interpreting and understanding the content of acquired data to generate appropriate information.

[0228] A "generative AI model" is a trained algorithm that uses artificial intelligence technology to analyze data and generate appropriate suggestions and answers.

[0229] Honorific language is the use of words to show respect to the other person, and is part of language usage.

[0230] "Word suggestion" is the act of selecting appropriate words and providing them to the user.

[0231] A "smart wearable device" is a portable electronic device that is worn by a user.

[0232] "Question data" is data of statements or text that a user uses to ask for specific information.

[0233] "Visual feedback" is a method of providing information to users in a form that they can see.

[0234] "Auditory feedback" is a method of providing information to a user in a form that can be heard by ear.

[0235] "Navigation information" is information that provides directions and location information to a specific location.

[0236] "Feedback" is the act of returning information or results from a system to a user.

[0237] The present invention is a communication support system for improving the efficiency of customer service in brick-and-mortar stores and enhancing customer satisfaction. This system is composed of three main elements: a user, a terminal (a smart wearable device), and a server. Specific embodiments for implementing the present invention are described below.

[0238] Acquisition and analysis of audio data

[0239] The user wears a device such as smart glasses or a smartwatch and responds to customers. The device captures the user's voice data in real time and sends that data to a server. The server then analyzes the voice data using a generative AI model and generates appropriate honorific and word suggestions. The generated suggestions are then fed back to the user visually or audibly via the device.

[0240] Image data acquisition and analysis

[0241] The user acquires image data of the interior of a store and the products through the device. The device then sends the captured image data to the server. The server analyzes the image data and generates location descriptions and names of objects. For example, it can also provide navigation information such as "The conference room is further ahead on the right."

[0242] Real-time question answering

[0243] When a user receives a question from a customer, the device sends the question data as voice or text to the server. The server analyzes the question data using a generative AI model and generates the optimal answer. This answer is also fed back to the user via the device.

[0244] Specific examples from physical stores

[0245] Example 1: Customer Service

[0246] When a customer asks, "Tell me about this new product," the user's device sends the question to the server, which uses a generative AI model to generate an answer such as, "This new product uses the latest AI technology and has a user-friendly interface," and the device visually displays the answer to the user.

[0247] Example 2: In-store navigation

[0248] When a customer asks, "Where is the conference room?", the user captures an image of the surroundings through the smart glasses and sends it to the server. The server analyzes the image data and generates navigation information such as "The conference room is on the second floor, on the right," which is then visually presented to the user on the device.

[0249] Hardware and software used

[0250] Hardware:

[0251] Smart Glasses

[0252] Smartwatch

[0253] microphone

[0254] software:

[0255] Python

[0256] SpeechRecognition Library

[0257] PIL (Python Imaging Library)

[0258] Requests library

[0259] OpenAI API

[0260] Prompt Sentence Examples

[0261] Here are some examples of prompts you can give to your generative AI model:

[0262] Customer Question: What are the features of this product?

[0263] Generate the appropriate answer.

[0264]

[0265] Please describe the image. The image shows a conference room.

[0266] The present invention allows users in physical stores to smoothly handle customer inquiries, provide instant responses to questions, and provide in-store navigation, which is expected to improve the quality of service and increase customer satisfaction.

[0267] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0268] Step 1:

[0269] The user puts on the smart glasses or smartwatch and begins interacting with the customer.

[0270] Input: User's voice or image data

[0271] Output: Captured audio or image data

[0272] Specific operation: The device's microphone captures audio data in real time, and the camera captures image data.

[0273] Step 2:

[0274] The terminal transmits the acquired audio and image data to the server.

[0275] Input: Audio or image data captured by the device

[0276] Output: Audio or image data transferred to the server

[0277] Specific operation: The terminal transmits the collected audio or image data to the server via the network.

[0278] Step 3:

[0279] The server analyzes the audio data and generates appropriate honorifics and word suggestions.

[0280] Input: Audio data received by the server

[0281] Output: Generated honorifics and word suggestions

[0282] Specific operation: The server processes the speech data using a generative AI model to generate appropriate honorifics and word suggestions.

[0283] Step 4:

[0284] The server analyzes the image data and generates location descriptions and names of objects.

[0285] Input: Image data received by the server

[0286] Output: Generated place descriptions and object names

[0287] How it works: The server processes image data using a generative AI model to identify location descriptions and object names.

[0288] Step 5:

[0289] The server sends generated suggestions and information to the terminal, which then provides visual or auditory feedback to the user.

[0290] Input: Generated suggestions and information from the server

[0291] Output: Feedback information provided on the device by display or audio

[0292] Specific operation: The server sends the generated honorific suggestions and navigation information to the device, which then displays them on the screen or notifies the user by voice.

[0293] Step 6:

[0294] It analyzes question data obtained from users via smart wearable devices and generates optimal answers using a generative AI model.

[0295] Input: Question data from the user

[0296] Output: The generated answer

[0297] How it works: The device receives a question via voice or text and sends it to the server, which then uses a generative AI model to analyze the question data and generate an answer.

[0298] Step 7:

[0299] The server generates a response and sends it to the terminal, which then provides visual or auditory feedback to the user.

[0300] Input: Generated answer from the server

[0301] Output: Answer information displayed on the device or provided as audio

[0302] Specific operation: The answer generated by the server is sent to the terminal and presented in a format that is easy for the user to see or hear.

[0303] Step 8:

[0304] The server displays in-store navigation information based on the generated information.

[0305] Input: Location information and image data from the user

[0306] Output: Generated navigation information

[0307] Specific operation: The server generates in-store navigation information based on location information and image data and sends it to the device, which then visually guides the user along the way.

[0308] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0309] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this embodiment is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[0310] System configuration

[0311] User

[0312] This is the subject who uses the system to receive communication support. They wear devices such as smart glasses or smart watches.

[0313] Terminal

[0314] A device that captures audio and image data in real time. In this case, it refers to smart glasses and smart watches. The device is responsible for sending the captured data to a server.

[0315] server

[0316] This is a central processing unit that analyzes the received voice and image data. It uses generative AI models to analyze the data and generate appropriate suggestions, which are then sent back to the device.

[0317] Emotion Engine

[0318] This engine recognizes the user's emotions from their voice and image data. Based on the emotion data, the server generates appropriate suggestions.

[0319] Program processing

[0320] Acquisition and analysis of voice and emotion data

[0321] 1. The user puts on the smart glasses or smartwatch and turns on the device.

[0322] 2. The device activates its audio and image sensors, and the user initiates a conversation.

[0323] 3. The terminal acquires the user's voice and image data in real time.

[0324] 4. The terminal sends the acquired voice and image data to the server.

[0325] 5. The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[0326] 6. The server generates appropriate suggestions based on the emotion engine analysis, such as "Great work. What would you like to talk about?"

[0327] 7. The server sends the generated suggestions to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (sound presented through earphones).

[0328] Image data acquisition and analysis

[0329] 1. When a user wants to check a specific place or object, they capture its image data through smart glasses.

[0330] 2. The device sends the captured image data to the server.

[0331] 3. The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expression. For example, if a user looks anxious while searching for a conference room.

[0332] 4. The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[0333] 5. The server sends the generated information to the terminal, which provides it to the user.

[0334] Real-time question answering

[0335] 1. The user speaks or writes a specific question, for example, "What does this word mean?"

[0336] 2. The device sends the question data to the server.

[0337] 3. The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expressions. For example, if the user is in a hurry.

[0338] 4. The server generates a quick and detailed answer based on the urgency of the question, such as "The meaning of this word is 'Tango'."

[0339] 5. The server sends the generated answer to the device, which provides it to the user via voice or text.

[0340] Specific examples

[0341] Example 1: Use during a business meeting

[0342] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[0343] Example 2: Using navigation

[0344] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[0345] As described above, the present invention allows users to receive appropriate communication support in real time, thereby reducing difficulties in business and daily life. This system is available to a wide range of users, regardless of age or level of understanding of technology, and is useful to many people.

[0346] The processing flow will be explained below.

[0347] Step 1:

[0348] The user puts on the smart glasses or smart watch and turns on the device. The device starts up and connects to the network.

[0349] Step 2:

[0350] The device activates its audio and image sensors and prepares to capture audio and images in real time.

[0351] Step 3:

[0352] When a user starts a conversation and has difficulty choosing words or using honorific language, they speak out verbally.

[0353] Step 4:

[0354] The terminal acquires the user's voice data in real time.

[0355] Step 5:

[0356] The terminal transmits the acquired voice data to the server.

[0357] Step 6:

[0358] The server analyzes the received voice data using a generative AI model.

[0359] Step 7:

[0360] The emotion engine recognizes the user's emotion based on the voice data. For example, if the user says, "I'm sorry, I'm in trouble," the emotion engine recognizes the user's confusion.

[0361] Step 8:

[0362] The server generates appropriate honorifics and word suggestions based on the analysis results of the emotion engine. For example, it generates a suggestion such as, "You seem to be in trouble. What can I explain to you?"

[0363] Step 9:

[0364] The server sends the generated proposal to the terminal.

[0365] Step 10:

[0366] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[0367] Step 11:

[0368] When a user wants to view a particular object or location, they capture image data of it through the smart glasses.

[0369] Step 12:

[0370] The terminal transmits the captured image data to the server.

[0371] Step 13:

[0372] The server analyzes the image data.

[0373] Step 14:

[0374] The emotion engine recognizes the user's emotional state from their facial expressions, for example, when the user has an anxious expression.

[0375] Step 15:

[0376] The server generates information including a description of the location and the name of the object based on the image analysis results and the recognition results of the emotion engine. For example, it could generate information such as "This is the conference room. It's up ahead on your right," providing additional information to ease the user's anxiety.

[0377] Step 16:

[0378] The server transmits the generated information to the terminal.

[0379] Step 17:

[0380] The device provides navigation information to the user visually (on the smart glasses display).

[0381] Step 18:

[0382] The user speaks or texts a specific question.

[0383] Step 19:

[0384] The terminal transmits the question data to the server.

[0385] Step 20:

[0386] The server parses the question data.

[0387] Step 21:

[0388] The emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression when asking the question. For example, if it determines that the question is urgent,

[0389] Step 22:

[0390] The server generates a quick and detailed answer based on the evaluation of the emotion engine, for example, "The meaning of this word is 'Tango'."

[0391] Step 23:

[0392] The server generates a response and sends it to the terminal.

[0393] Step 24:

[0394] The device provides the user with a response by voice or text.

[0395] Step 25:

[0396] The user selects a subscription plan using a dedicated application.

[0397] Step 26:

[0398] The terminal transmits the selected plan information to the server.

[0399] Step 27:

[0400] The server enables certain features based on the plan selected by the user.

[0401] Step 28:

[0402] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[0403] Through this process, the system enables users to receive appropriate communication assistance in real time, easing difficulties in business and daily life.

[0404] Example 2

[0405] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0406] Conventional communication support systems have had difficulty in providing real-time support by fully utilizing the user's voice data and image data. In particular, they lacked the ability to provide appropriate suggestions and information that took the user's emotions into consideration, making it difficult to improve the user experience. They also lacked the ability to provide prompt and appropriate answers to user questions. It is necessary to provide a system that can solve these problems and enable users to receive better support.

[0407] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0408] In this invention, the server includes means for acquiring and processing a user's voice data and image data in real time, means for transmitting the voice data and image data to the server, means for analyzing the voice data and image data by the server and recognizing emotions, means for analyzing the voice data using a generative AI model to generate appropriate suggestions, and means for feeding back the generated suggestions to the user. This allows the user to receive appropriate suggestions and information that take emotions into consideration in real time, significantly improving the communication support experience. It also enables the provision of quick and appropriate answers to questions.

[0409] A "user" is an entity that utilizes the communication support system to provide voice and image data and receive support.

[0410] A "terminal" is a device that captures audio and image data in real time and transmits it to a server, and examples of this include smart glasses and smart watches.

[0411] The "server" is a central processing unit that analyzes the received audio and image data, generates appropriate suggestions, and sends them to the terminal.

[0412] The "emotion engine" is an analysis engine for recognizing the user's emotions from the user's voice data and image data.

[0413] A "generative AI model" is an artificial intelligence model that runs on a server and analyzes voice and image data to generate appropriate suggestions and answers.

[0414] "Voice data" refers to data that is a digital recording of a user's speech.

[0415] "Image data" is digital data that captures the user's face and environment with a camera.

[0416] "Analysis" is the process of extracting information from the acquired voice data and image data and understanding their meaning.

[0417] "Suggestions" are advice or information for the user that are generated by the server based on the results of the analysis.

[0418] "Question data" is voice or text data containing a specific question from the user.

[0419] An "answer" is a solution generated by the server by analyzing question data, and is information provided to the user.

[0420] "Feedback" is the act of providing suggestions or answers to a user, and can be done visually or audibly.

[0421] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this invention is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[0422] First, the user puts on a device such as smart glasses or a smart watch and turns it on. The device activates its audio and image sensors, and the user begins a conversation. The device captures the user's audio and image data in real time and sends the acquired data to a server. The device can be a common wearable device such as smart glasses or a smart watch. These devices communicate with the server via Wi-Fi or Bluetooth.

[0423] Next, the server analyzes the received voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to talk about," the emotion engine recognizes that the user is confused. Based on the emotion engine's analysis results, the server generates an appropriate suggestion, such as, "Thank you for your hard work. What would you like to talk about?" The server then sends the generated suggestion to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (audio provided through earphones).

[0424] Furthermore, if a user wants to check a specific location or object, they can capture image data of that location through the smart glasses. The device sends the captured image data to a server, which then analyzes the image data. This analysis includes cases where the user has an anxious expression, such as when searching for a conference room. The server generates navigation information such as "The conference room is further ahead on your right," and provides additional information that takes the user's emotions into consideration. The server then sends the generated information to the device, which then provides it to the user.

[0425] In addition, a real-time question-and-answer function is also provided. The user inputs a specific question by voice or text, and the device sends the question data to the server. The server analyzes the question data, and an emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression. For example, if the user is in a hurry, the server will take the urgency into account and generate a quick and detailed answer such as, "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[0426] Specific examples

[0427] Example 1: Use during a business meeting

[0428] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[0429] Example 2: Using navigation

[0430] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[0431] Prompt Sentence Examples

[0432] Below are some examples of prompts to input to the generative AI model:

[0433] "Please analyze the audio data that the user is confused about."

[0434] "Analyze image data showing a user's anxious expression."

[0435] "Generate quick answers to user questions"

[0436] As a result, users can receive appropriate communication support in real time, and difficulties in business and daily life can be alleviated.

[0437] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0438] Step 1:

[0439] The user puts on the smart glasses or smartwatch and turns on the device.

[0440] Input: A user presses the power button on the device.

[0441] Operation: The device boots up and performs self-tests and initialization.

[0442] Output: The device displays a status of ready.

[0443] Step 2:

[0444] The device activates its audio and image sensors, and the user initiates a conversation.

[0445] Input: Device readiness status and user voice input.

[0446] Action: The device activates the audio sensors and camera and begins capturing data.

[0447] Output: Audio and image data captured by the device.

[0448] Step 3:

[0449] The terminal acquires the user's voice and image data in real time.

[0450] Input: User's speech and facial expressions.

[0451] How it works: The audio sensor converts sound into digital data, and the camera captures images.

[0452] Output: Acquired audio and image data.

[0453] Step 4:

[0454] The terminal transmits the acquired voice data and image data to the server.

[0455] Input: Captured audio and image data.

[0456] How it works: Your device sends data to a server via Wi-Fi or Bluetooth.

[0457] Output: Audio and image data sent to the server.

[0458] Step 5:

[0459] The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions.

[0460] Input: Transmitted audio and image data.

[0461] How it works: The server inputs voice data into a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[0462] Output: Parsed sentiment data and appropriate suggestions.

[0463] Step 6:

[0464] The server generates appropriate suggestions based on the analysis results of the emotion engine, such as "Good work. What would you like to talk about?"

[0465] Input: Parsed emotion data and the analysis results of the generative AI model.

[0466] How it works: The server generates appropriate suggestions and converts them into text or audio format.

[0467] Output: The generated proposals.

[0468] Step 7:

[0469] The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[0470] Input: Generated proposals.

[0471] How it works: The server sends the suggestions to the device, which then provides feedback to the user, for example, on the smart glasses display or as audio through earphones.

[0472] Output: Feedback provided to the user.

[0473] Step 8:

[0474] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[0475] Input: The situation in which you are looking at the place or thing you want to check.

[0476] How it works: The user operates the smart glasses and the camera captures image data.

[0477] Output: The captured image data.

[0478] Step 9:

[0479] The terminal transmits the captured image data to the server.

[0480] Input: The captured image data.

[0481] Operation: The device compresses the image data and sends it to the server.

[0482] Output: Image data sent to the server.

[0483] Step 10:

[0484] The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expressions.

[0485] Input: The submitted image data.

[0486] How it works: The server analyzes the image data, and the emotion engine recognizes facial expressions. For example, it analyzes the facial expressions of a user searching for a meeting room.

[0487] Output: Parsed emotion data and pertinent information.

[0488] Step 11:

[0489] The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[0490] Input: Parsed emotion data and user location information.

[0491] How it works: The server generates navigation information and emotion-sensitive messages.

[0492] Output: Generated navigation information and additional information.

[0493] Step 12:

[0494] The server transmits the generated information to the terminal, which then provides it to the user.

[0495] Input: Generated navigation information and additional information.

[0496] How it works: The server sends information to the device, which then provides it to the user, for example by displaying it on smart glasses or providing audio guidance.

[0497] Output: Information provided to the user.

[0498] Step 13:

[0499] The user speaks or texts a specific question.

[0500] Input: The user's question.

[0501] How it works: A user types a question using the microphone or text input function on their smart glasses.

[0502] Output: The input question data.

[0503] Step 14:

[0504] The terminal transmits the question data to the server.

[0505] Input: The entered question data.

[0506] Operation: The device sends the query data to the server.

[0507] Output: The query data sent to the server.

[0508] Step 15:

[0509] The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression.

[0510] Input: Submitted question data and user sentiment data.

[0511] How it works: The server analyzes the question data, and the emotion engine evaluates the urgency and importance.

[0512] Output: Assessed urgency and importance.

[0513] Step 16:

[0514] The server takes into account the urgency and generates a quick and detailed response.

[0515] Input: Assessed urgency and question data.

[0516] How it works: The server uses a generative AI model to generate an answer, such as "The meaning of this word is 'Tango'."

[0517] Output: The generated answer.

[0518] Step 17:

[0519] The server generates a response and sends it to the terminal, which provides it to the user in voice or text.

[0520] Input: The generated answer.

[0521] Operation: The server sends the answer data to the terminal, which then provides it to the user via voice or text.

[0522] Output: The answer provided to the user.

[0523] (Application example 2)

[0524] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0525] Current communication support systems can make appropriate suggestions using the user's voice and image data, but they cannot generate responses based on the user's emotions, making it difficult to provide prompt and appropriate support.In addition, in certain environments such as factories, there are no systems that can properly analyze and respond to workers' difficulties and stress, making it difficult to improve productivity.

[0526] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, and means including an emotion engine for analyzing the user's emotion data and generating responses. This makes it possible to provide the user with appropriate suggestions and responses based on the emotion data, thereby achieving more effective communication support.

[0527] A "user" is an entity that uses the system to receive communication support.

[0528] "Voice data" refers to data that records the voice uttered by the user.

[0529] A "server" is a central processing unit that analyzes the data received and generates the necessary suggestions and responses.

[0530] "Emotion data" refers to data that indicates emotional information extracted from the user's voice or image.

[0531] An "emotion engine" is a system component that analyzes emotion data and recognizes the user's emotions.

[0532] A "generative AI model" is an artificial intelligence model that generates appropriate responses and suggestions based on input data.

[0533] "Suggestion" refers to the generation of useful information or advice for the user.

[0534] "Feedback" is the act of providing server-generated suggestions and responses to the user.

[0535] "Question data" is data including the content of a question provided by a user.

[0536] This invention provides a system for an emotionally responsive communication robot that supports workers in factories. Specifically, the robot acquires the worker's voice data and image data in real time, sends this data to a server for analysis, and generates appropriate suggestions and responses based on the worker's emotions. This reduces worker stress and helps improve productivity.

[0537] Hardware and software used

[0538] Hardware: A robotic device equipped with audio and image sensors, a server as the central processing unit.

[0539] Software: Speech analysis software (Google Cloud Speech-to-Text API), image analysis software (OpenCV), emotion engine (Microsoft Azure Emotion API), generative AI model (OpenAI GPT-4).

[0540] Program processing

[0541] 1. Acquisition and analysis of voice and emotion data

[0542] The worker begins working in front of the robot.

[0543] The robot activates audio and image sensors to capture the worker's voice and facial expressions in real time.

[0544] The robot transmits the acquired voice data and image data to the server.

[0545] The server analyzes the voice data using the Google Cloud Speech-to-Text API and the emotion data using the Microsoft Azure Emotion API.

[0546] Based on the emotional data, suggestions tailored to the worker's situation are generated using a generative AI model (OpenAI GPT-4).

[0547] The server sends the generated suggestions back to the robot, which then provides the suggestions to the worker via voice or display.

[0548] 2. Image Data Acquisition and Analysis

[0549] If a worker is having difficulty with a particular task, the robot will capture that situation.

[0550] The robot transmits the captured image data to a server.

[0551] The server uses OpenCV to analyze the image data, and the emotion engine recognizes the emotions from the worker's facial expressions.

[0552] Appropriate suggestions based on the worker's emotions are generated by a generative AI model, and then communicated by the robot.

[0553] Specific examples

[0554] 1. Support for workers

[0555] When a worker is performing a difficult task, the robot detects the "difficulty" from the worker's voice and facial expression.

[0556] The robot will make suggestions such as, "Thank you for your hard work. Is there anything I can help you with?"

[0557] This allows workers to receive appropriate support and carry out their work efficiently.

[0558] 2. Suggest short breaks

[0559] If a worker looks tired, the emotion engine will recognize this.

[0560] An example of a prompt from a generative AI model: "You seem to be tired. How about taking a short break?" Based on this, the model suggests, "How about taking a short break?"

[0561] This reduces worker fatigue and improves productivity.

[0562] In this way, communication support can be realized by providing appropriate suggestions and responses based on the user's emotions. This system is also very useful for supporting workers in factories and other specific environments.

[0563] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0564] Step 1:

[0565] When the user starts working in front of the robot, the robot activates the audio and image sensors. The audio sensor then captures the user's voice data, and the image sensor captures the user's facial expressions. Both data (audio and image data) are acquired in real time.

[0566] Input: User's voice and image

[0567] Output: Captured audio and image data

[0568] Step 2:

[0569] The device sends the acquired audio and image data to the server, where it is ready to be analyzed.

[0570] Input: Acquired audio and image data

[0571] Output: Audio and image data sent to the server

[0572] Step 3:

[0573] The server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then used for analysis by the emotion engine. It also uses OpenCV to extract facial expression data from the image data.

[0574] Input: Audio and image data sent to the server

[0575] Output: Analyzed voice data (text data) and facial expression data

[0576] Step 4:

[0577] The emotion engine (Microsoft Azure Emotion API) analyzes the user's emotions from the voice and facial expression data. The emotion data indicates the user's current emotional state (e.g., confusion, fatigue, tension, etc.).

[0578] Input: Analyzed voice data and facial expression data

[0579] Output: User emotion data

[0580] Step 5:

[0581] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate suggestions and responses based on the emotion data. Specific prompts are provided to the generative AI model using the emotion data as input.

[0582] Input: User emotion data

[0583] Output: The generated suggestions and responses

[0584] Step 6:

[0585] The server sends the generated suggestions and responses to the robot, which then provides the suggestions and responses to the user through voice or a display. For example, a message such as "Good work! Is there anything I can help you with?" may be displayed or spoken.

[0586] Input: Generated suggestions and responses

[0587] Output: The suggestions or responses provided to the user

[0588] Step 7:

[0589] The user receives suggestions and responses from the robot and, if necessary, further communicates with the robot. The user's reactions and additional audio and video data are processed again starting from step 1.

[0590] Input: User reactions to the suggestions and responses provided

[0591] Output: New audio and image data for the next cycle

[0592] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0593] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0594] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0595] [Second embodiment]

[0596] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0597] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0598] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0599] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0600] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0601] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0602] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0603] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0604] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0605] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0606] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0607] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0608] This invention is a communication support system that uses AI as its underlying technology, and analyzes the user's voice and image data in real time to suggest appropriate honorific expressions and words, as well as provide answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[0609] System configuration

[0610] User

[0611] This is the entity that uses the system to receive communication support. They use devices such as smart glasses or smart watches.

[0612] Terminal

[0613] A device that captures audio and image data in real time. This includes smart glasses and smart watches. The device transmits the captured data to a server.

[0614] server

[0615] It is a central processing unit that analyzes the received voice and image data. The server uses a generative AI model to analyze the data, generate appropriate suggestions and answers, and send them back to the device.

[0616] Program processing

[0617] Acquisition and analysis of audio data

[0618] 1. The user puts on the smart glasses or smartwatch and starts a conversation.

[0619] 2. The device captures the user's voice data in real time and sends the data to the server.

[0620] 3. The server analyzes the received voice data using a generative AI model. For example, if a user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates the suggestion "It's an honor to meet you."

[0621] 4. The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[0622] Image data acquisition and analysis

[0623] 1. When a user wants to see a particular place or object, they capture an image of it through the smart glasses.

[0624] 2. The device sends the captured image data to the server.

[0625] 3. The server analyzes the received image data and generates a description of the location or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on the right."

[0626] 4. The server sends the generated information to the terminal, which provides it to the user.

[0627] Real-time question answering

[0628] 1. A user asks a specific question via voice or text, for example, "What does this word mean?"

[0629] 2. The device sends the question data to the server.

[0630] 3. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'."

[0631] 4. The server generates a response and sends it to the device, which provides it to the user via voice or text.

[0632] Specific examples

[0633] Example 1: Use during a business meeting

[0634] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[0635] Example 2: Using navigation

[0636] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[0637] As described above, the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. This system is usable by a wide range of users regardless of age or level of understanding of technology, making it useful to many people.

[0638] The processing flow will be explained below.

[0639] Step 1:

[0640] The user puts on the smart glasses or smartwatch and turns on the device, which starts up the device and connects it to the network.

[0641] Step 2:

[0642] The device activates its audio and image sensors and prepares to capture data in real time.

[0643] Step 3:

[0644] When a user starts a conversation and is unsure about the meaning of a word or how to use honorific language, they speak out verbally.

[0645] Step 4:

[0646] The terminal acquires the user's voice data in real time.

[0647] Step 5:

[0648] The terminal transmits the acquired voice data to the server.

[0649] Step 6:

[0650] The server analyzes the received voice data using a generative AI model. For example, if a user says, "Oh, sorry. What was it?", the server analyzes the voice data and generates an appropriate polite expression such as, "Thank you for your hard work. What are you looking for?"

[0651] Step 7:

[0652] The server sends the analysis results to the device.

[0653] Step 8:

[0654] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[0655] Step 9:

[0656] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[0657] Step 10:

[0658] The terminal transmits the captured image data to the server.

[0659] Step 11:

[0660] The server analyzes the received image data and generates information such as "The conference room is further ahead on the right."

[0661] Step 12:

[0662] The server transmits the generated information to the terminal.

[0663] Step 13:

[0664] The device provides navigation information to the user visually (on the smart glasses display).

[0665] Step 14:

[0666] The user speaks or texts a specific question, for example, "What does this word mean?"

[0667] Step 15:

[0668] The terminal transmits the question data to the server.

[0669] Step 16:

[0670] The server analyzes the question data and generates an answer such as "The meaning of this word is 'Tango'."

[0671] Step 17:

[0672] The server sends the generated response to the terminal.

[0673] Step 18:

[0674] The device provides the user with a response by voice or text.

[0675] Step 19:

[0676] The user selects a subscription plan using a dedicated application.

[0677] Step 20:

[0678] The terminal transmits the selected plan information to the server.

[0679] Step 21:

[0680] The server enables certain features based on the plan selected by the user.

[0681] Step 22:

[0682] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[0683] This allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life.

[0684] Example 1

[0685] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0686] In modern society, communication is extremely important in business and everyday life. However, many people have difficulty using appropriate honorifics and vocabulary, and recognizing places and objects, as well as answering questions instantly, are now essential needs. Conventional communication support systems lack the ability to efficiently solve these problems in real time. In particular, more advanced technology is needed to provide appropriate content suggestions and rapid analytical feedback in a variety of situations.

[0687] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0688] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to a central processing unit, means for analyzing the voice data by the central processing unit and generating suggestions for appropriate honorific expressions and words, means for feeding back the generated suggestions to the user, and a generative AI model for analyzing the voice data. As a result, the user can receive suggestions for appropriate honorific expressions and words in real time, improving the quality of communication and reducing difficulties in business and daily life.

[0689] A "user" is an entity that uses the system to receive communication support.

[0690] "Voice data" refers to data in which the user's voice is recorded in digital form.

[0691] "Real-time" means that processing occurs immediately and on the spot with minimal delay.

[0692] "Means of acquisition" refers to devices and functions for collecting audio data and image data.

[0693] A "central processing unit" is a server or computer system that performs major processing such as data analysis.

[0694] "Means of transmission" refers to the communications means or technology used to transfer acquired data to a specific destination.

[0695] "Means for analyzing" refers to software or algorithms used to analyze received data.

[0696] "Means for generating suggestions" refers to a function that generates appropriate information and advice to be provided to the user based on the analyzed data.

[0697] "Feedback means" refers to an output device or method for notifying the user of the generated suggestions.

[0698] "Image data" is data that records a user's visual information in digital form.

[0699] "Location description" refers to content that provides information or directions to a specific location.

[0700] "Name of object" refers to the name of the object contained in the image data.

[0701] "Question data" refers to data that indicates the content of a question that a user asks the system.

[0702] "Means for generating answers" refers to algorithms or software for generating optimal answers based on question data.

[0703] A "generative AI model" is a model or system that uses artificial intelligence to analyze data and generate suggestions or answers.

[0704] This invention is a communication support system that uses a generative AI model as its underlying technology. It analyzes the user's voice and image data in real time, and suggests appropriate honorific expressions and words, as well as providing answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[0705] Acquisition and analysis of audio data

[0706] A user puts on smart glasses or a smartwatch and starts a conversation. The device captures the user's voice data in real time and sends it to a server. The server analyzes the received voice data using a generative AI model (e.g., OpenAI's GPT-4). For example, if the user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates a suggestion: "It's an honor to meet you." The server sends the suggestion to the device, which then provides visual or auditory feedback to the user.

[0707] Image data acquisition and analysis

[0708] When a user wants to check a specific place or object, they capture an image of it through the smart glasses. The device sends the captured image data to a server. The server analyzes the received image data and generates a description of the place or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on your right." The server sends the generated information to the device, which then provides it to the user.

[0709] Real-time question answering

[0710] The user asks a specific question by voice or text. For example, "What does this word mean?", the device sends the question data to the server. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[0711] Specific examples

[0712] Example 1: Use during a business meeting

[0713] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[0714] Example 2: Using navigation

[0715] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[0716] Prompt Sentence Examples

[0717] Use the following prompt:

[0718] User Input: "Nice to meet you."

[0719] AI Model: "It's a pleasure to meet you."

[0720] User Input: "What does this word mean?"

[0721] AI Model: "This word means 'Tango'"

[0722] In this way, the system of the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. Furthermore, by using a generative AI model, it is possible to respond quickly and accurately to user needs, making the system useful for many people, regardless of age or level of technological literacy.

[0723] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0724] Acquisition and analysis of audio data

[0725] Step 1: The user puts on the smart glasses or smartwatch and starts a conversation.

[0726] Input: User's voice

[0727] Action: The user speaks.

[0728] Output: Audio signal

[0729] Step 2: The terminal acquires the user's voice data in real time.

[0730] Input: Audio signal

[0731] How it works: The device's microphone picks up audio signals and converts them into digital audio data.

[0732] Output: Digital audio data

[0733] Step 3: The device sends the acquired voice data to the server.

[0734] Input: Digital audio data

[0735] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[0736] Output: Compressed and encrypted audio data

[0737] Step 4: The server analyzes the received voice data using the generative AI model.

[0738] Input: Compressed and encrypted audio data

[0739] How it works: The server decodes the audio data and uses a generative AI model (e.g., GPT-4) to convert the audio to text and perform analysis.

[0740] Output: Analysis results (suggestions for appropriate honorifics and words)

[0741] Step 5: The server sends the generated proposal to the device.

[0742] Input: Analysis results (suggestions for appropriate honorifics and words)

[0743] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[0744] Output: Compressed and encrypted proposal data

[0745] Step 6: The device provides visual or auditory feedback to the user.

[0746] Input: Compressed and encrypted proposal data

[0747] How it works: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[0748] Output: User feedback (visual / auditory information)

[0749] Image data acquisition and analysis

[0750] Step 1: When a user wants to see a specific place or object, they capture an image of it through their smart glasses.

[0751] Input: Visual information

[0752] Action: The user operates the smartglasses camera to capture an image.

[0753] Output: Image data

[0754] Step 2: The device sends the captured image data to the server.

[0755] Input: Image data

[0756] Operation: The terminal's communication module compresses and encrypts the image data and sends it to the server.

[0757] Output: Compressed and encrypted image data

[0758] Step 3: The server analyzes the received image data.

[0759] Input: Compressed and encrypted image data

[0760] How it works: The server decodes the data and uses image recognition software to analyze the image (e.g., recognize location descriptions and object names).

[0761] Output: Analysis results (location descriptions and names of things)

[0762] Step 4: The server sends the generated information to the terminal.

[0763] Input: Analysis results (location description and object names)

[0764] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[0765] Output: Compressed and encrypted information data

[0766] Step 5: The terminal provides the user with information visually or audibly.

[0767] Input: Compressed and encrypted information data

[0768] How it works: The device decodes the data and displays it on the screen or speaks the information over the speaker.

[0769] Output: Provide information to the user (visual / auditory information)

[0770] Real-time question answering

[0771] Step 1: The user asks a specific question via voice or text.

[0772] Input: Voice or text question

[0773] How it works: If the user asks a question by voice, the device's microphone picks up the sound, and if the user asks a question by text, the device's text input function is used.

[0774] Output: Digital audio data or text data

[0775] Step 2: The device sends the query data to the server.

[0776] Input: Digital audio or text data

[0777] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[0778] Output: Compressed and encrypted query data

[0779] Step 3: The server analyzes the question data and generates the best answer.

[0780] Input: Compressed and encrypted query data

[0781] How it works: The server decodes the data, uses a generative AI model to analyze the intent of the question, and generates an answer.

[0782] Output: Analysis results (best answer)

[0783] Step 4: The server sends the generated response to the terminal.

[0784] Input: Analysis results (best answer)

[0785] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[0786] Output: Compressed and encrypted answer data

[0787] Step 5: The device provides the user with a voice or text message.

[0788] Input: Compressed and encrypted answer data

[0789] What it does: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[0790] Output: Provides answers to users (visual / auditory information)

[0791] As described above, the system achieves efficient communication support for users through specific operations and data input / output at each step.

[0792] (Application example 1)

[0793] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0794] Conventional communication support systems allow users to receive suggestions for appropriate honorific expressions and words, but they have difficulty answering customer questions in real time or providing in-store navigation information when dealing with customers in brick-and-mortar stores. They also lack the ability to provide personalized services based on customer facial recognition or past purchase history. The present invention provides a new communication support system that streamlines customer service in brick-and-mortar stores and improves customer satisfaction.

[0795] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0796] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, means for analyzing question data acquired from the user via the smart wearable device and generating an optimal answer using a generative AI model, means for visually or audibly feeding back the generated answer, and means for displaying in-store navigation information based on the generated information, thereby enabling immediate responses to customer questions in physical stores, in-store navigation, and personalized services.

[0797] "Voice data" is information about sounds generated by the user's vocalizations.

[0798] A "server" is a central processing unit that processes and provides various data via a network.

[0799] "Analysis" is the act of interpreting and understanding the content of acquired data to generate appropriate information.

[0800] A "generative AI model" is a trained algorithm that uses artificial intelligence technology to analyze data and generate appropriate suggestions and answers.

[0801] Honorific language is the use of words to show respect to the other person, and is part of language usage.

[0802] "Word suggestion" is the act of selecting appropriate words and providing them to the user.

[0803] A "smart wearable device" is a portable electronic device that is worn by a user.

[0804] "Question data" is data of statements or text that a user uses to ask for specific information.

[0805] "Visual feedback" is a method of providing information to users in a form that they can see.

[0806] "Auditory feedback" is a method of providing information to a user in a form that can be heard by ear.

[0807] "Navigation information" is information that provides directions and location information to a specific location.

[0808] "Feedback" is the act of returning information or results from a system to a user.

[0809] The present invention is a communication support system for improving the efficiency of customer service in brick-and-mortar stores and enhancing customer satisfaction. This system is composed of three main elements: a user, a terminal (a smart wearable device), and a server. Specific embodiments for implementing the present invention are described below.

[0810] Acquisition and analysis of audio data

[0811] The user wears a device such as smart glasses or a smartwatch and responds to customers. The device captures the user's voice data in real time and sends that data to a server. The server then analyzes the voice data using a generative AI model and generates appropriate honorific and word suggestions. The generated suggestions are then fed back to the user visually or audibly via the device.

[0812] Image data acquisition and analysis

[0813] The user acquires image data of the interior of a store and the products through the device. The device then sends the captured image data to the server. The server analyzes the image data and generates location descriptions and names of objects. For example, it can also provide navigation information such as "The conference room is further ahead on the right."

[0814] Real-time question answering

[0815] When a user receives a question from a customer, the device sends the question data as voice or text to the server. The server analyzes the question data using a generative AI model and generates the optimal answer. This answer is also fed back to the user via the device.

[0816] Specific examples from physical stores

[0817] Example 1: Customer Service

[0818] When a customer asks, "Tell me about this new product," the user's device sends the question to the server, which uses a generative AI model to generate an answer such as, "This new product uses the latest AI technology and has a user-friendly interface," and the device visually displays the answer to the user.

[0819] Example 2: In-store navigation

[0820] When a customer asks, "Where is the conference room?", the user captures an image of the surroundings through the smart glasses and sends it to the server. The server analyzes the image data and generates navigation information such as "The conference room is on the second floor, on the right," which is then visually presented to the user on the device.

[0821] Hardware and software used

[0822] Hardware:

[0823] Smart Glasses

[0824] Smartwatch

[0825] microphone

[0826] software:

[0827] Python

[0828] SpeechRecognition Library

[0829] PIL (Python Imaging Library)

[0830] Requests library

[0831] OpenAI API

[0832] Prompt Sentence Examples

[0833] Here are some examples of prompts you can give to your generative AI model:

[0834] Customer Question: What are the features of this product?

[0835] Generate the appropriate answer.

[0836]

[0837] Please describe the image. The image shows a conference room.

[0838] The present invention allows users in physical stores to smoothly handle customer inquiries, provide instant responses to questions, and provide in-store navigation, which is expected to improve the quality of service and increase customer satisfaction.

[0839] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0840] Step 1:

[0841] The user puts on the smart glasses or smartwatch and begins interacting with the customer.

[0842] Input: User's voice or image data

[0843] Output: Captured audio or image data

[0844] Specific operation: The device's microphone captures audio data in real time, and the camera captures image data.

[0845] Step 2:

[0846] The terminal transmits the acquired audio and image data to the server.

[0847] Input: Audio or image data captured by the device

[0848] Output: Audio or image data transferred to the server

[0849] Specific operation: The terminal transmits the collected audio or image data to the server via the network.

[0850] Step 3:

[0851] The server analyzes the audio data and generates appropriate honorifics and word suggestions.

[0852] Input: Audio data received by the server

[0853] Output: Generated honorifics and word suggestions

[0854] Specific operation: The server processes the speech data using a generative AI model to generate appropriate honorifics and word suggestions.

[0855] Step 4:

[0856] The server analyzes the image data and generates location descriptions and names of objects.

[0857] Input: Image data received by the server

[0858] Output: Generated place descriptions and object names

[0859] How it works: The server processes image data using a generative AI model to identify location descriptions and object names.

[0860] Step 5:

[0861] The server sends generated suggestions and information to the terminal, which then provides visual or auditory feedback to the user.

[0862] Input: Generated suggestions and information from the server

[0863] Output: Feedback information provided on the device by display or audio

[0864] Specific operation: The server sends the generated honorific suggestions and navigation information to the device, which then displays them on the screen or notifies the user by voice.

[0865] Step 6:

[0866] It analyzes question data obtained from users via smart wearable devices and generates optimal answers using a generative AI model.

[0867] Input: Question data from the user

[0868] Output: The generated answer

[0869] How it works: The device receives a question via voice or text and sends it to the server, which then uses a generative AI model to analyze the question data and generate an answer.

[0870] Step 7:

[0871] The server generates a response and sends it to the terminal, which then provides visual or auditory feedback to the user.

[0872] Input: Generated answer from the server

[0873] Output: Answer information displayed on the device or provided as audio

[0874] Specific operation: The answer generated by the server is sent to the terminal and presented in a format that is easy for the user to see or hear.

[0875] Step 8:

[0876] The server displays in-store navigation information based on the generated information.

[0877] Input: Location information and image data from the user

[0878] Output: Generated navigation information

[0879] Specific operation: The server generates in-store navigation information based on location information and image data and sends it to the device, which then visually guides the user along the way.

[0880] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0881] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this embodiment is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[0882] System configuration

[0883] User

[0884] This is the subject who uses the system to receive communication support. They wear devices such as smart glasses or smart watches.

[0885] Terminal

[0886] A device that captures audio and image data in real time. In this case, it refers to smart glasses and smart watches. The device is responsible for sending the captured data to a server.

[0887] server

[0888] This is a central processing unit that analyzes the received voice and image data. It uses generative AI models to analyze the data and generate appropriate suggestions, which are then sent back to the device.

[0889] Emotion Engine

[0890] This engine recognizes the user's emotions from their voice and image data. Based on the emotion data, the server generates appropriate suggestions.

[0891] Program processing

[0892] Acquisition and analysis of voice and emotion data

[0893] 1. The user puts on the smart glasses or smartwatch and turns on the device.

[0894] 2. The device activates its audio and image sensors, and the user initiates a conversation.

[0895] 3. The terminal acquires the user's voice and image data in real time.

[0896] 4. The terminal sends the acquired voice and image data to the server.

[0897] 5. The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[0898] 6. The server generates appropriate suggestions based on the emotion engine analysis, such as "Great work. What would you like to talk about?"

[0899] 7. The server sends the generated suggestions to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (sound presented through earphones).

[0900] Image data acquisition and analysis

[0901] 1. When a user wants to check a specific place or object, they capture its image data through smart glasses.

[0902] 2. The device sends the captured image data to the server.

[0903] 3. The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expression. For example, if a user looks anxious while searching for a conference room.

[0904] 4. The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[0905] 5. The server sends the generated information to the terminal, which provides it to the user.

[0906] Real-time question answering

[0907] 1. The user speaks or writes a specific question, for example, "What does this word mean?"

[0908] 2. The device sends the question data to the server.

[0909] 3. The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expressions. For example, if the user is in a hurry.

[0910] 4. The server generates a quick and detailed answer based on the urgency of the question, such as "The meaning of this word is 'Tango'."

[0911] 5. The server sends the generated answer to the device, which provides it to the user via voice or text.

[0912] Specific examples

[0913] Example 1: Use during a business meeting

[0914] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[0915] Example 2: Using navigation

[0916] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[0917] As described above, the present invention allows users to receive appropriate communication support in real time, thereby reducing difficulties in business and daily life. This system is available to a wide range of users, regardless of age or level of understanding of technology, and is useful to many people.

[0918] The processing flow will be explained below.

[0919] Step 1:

[0920] The user puts on the smart glasses or smart watch and turns on the device. The device starts up and connects to the network.

[0921] Step 2:

[0922] The device activates its audio and image sensors and prepares to capture audio and images in real time.

[0923] Step 3:

[0924] When a user starts a conversation and has difficulty choosing words or using honorific language, they speak out verbally.

[0925] Step 4:

[0926] The terminal acquires the user's voice data in real time.

[0927] Step 5:

[0928] The terminal transmits the acquired voice data to the server.

[0929] Step 6:

[0930] The server analyzes the received voice data using a generative AI model.

[0931] Step 7:

[0932] The emotion engine recognizes the user's emotion based on the voice data. For example, if the user says, "I'm sorry, I'm in trouble," the emotion engine recognizes the user's confusion.

[0933] Step 8:

[0934] The server generates appropriate honorifics and word suggestions based on the analysis results of the emotion engine. For example, it generates a suggestion such as, "You seem to be in trouble. What can I explain to you?"

[0935] Step 9:

[0936] The server sends the generated proposal to the terminal.

[0937] Step 10:

[0938] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[0939] Step 11:

[0940] When a user wants to view a particular object or location, they capture image data of it through the smart glasses.

[0941] Step 12:

[0942] The terminal transmits the captured image data to the server.

[0943] Step 13:

[0944] The server analyzes the image data.

[0945] Step 14:

[0946] The emotion engine recognizes the user's emotional state from their facial expressions, for example, when the user has an anxious expression.

[0947] Step 15:

[0948] The server generates information including a description of the location and the name of the object based on the image analysis results and the recognition results of the emotion engine. For example, it could generate information such as "This is the conference room. It's up ahead on your right," providing additional information to ease the user's anxiety.

[0949] Step 16:

[0950] The server transmits the generated information to the terminal.

[0951] Step 17:

[0952] The device provides navigation information to the user visually (on the smart glasses display).

[0953] Step 18:

[0954] The user speaks or texts a specific question.

[0955] Step 19:

[0956] The terminal transmits the question data to the server.

[0957] Step 20:

[0958] The server parses the question data.

[0959] Step 21:

[0960] The emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression when asking the question. For example, if it determines that the question is urgent,

[0961] Step 22:

[0962] The server generates a quick and detailed answer based on the evaluation of the emotion engine, for example, "The meaning of this word is 'Tango'."

[0963] Step 23:

[0964] The server generates a response and sends it to the terminal.

[0965] Step 24:

[0966] The device provides the user with a response by voice or text.

[0967] Step 25:

[0968] The user selects a subscription plan using a dedicated application.

[0969] Step 26:

[0970] The terminal transmits the selected plan information to the server.

[0971] Step 27:

[0972] The server enables certain features based on the plan selected by the user.

[0973] Step 28:

[0974] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[0975] Through this process, the system enables users to receive appropriate communication assistance in real time, easing difficulties in business and daily life.

[0976] Example 2

[0977] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0978] Conventional communication support systems have had difficulty in providing real-time support by fully utilizing the user's voice data and image data. In particular, they lacked the ability to provide appropriate suggestions and information that took the user's emotions into consideration, making it difficult to improve the user experience. They also lacked the ability to provide prompt and appropriate answers to user questions. It is necessary to provide a system that can solve these problems and enable users to receive better support.

[0979] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0980] In this invention, the server includes means for acquiring and processing a user's voice data and image data in real time, means for transmitting the voice data and image data to the server, means for analyzing the voice data and image data by the server and recognizing emotions, means for analyzing the voice data using a generative AI model to generate appropriate suggestions, and means for feeding back the generated suggestions to the user. This allows the user to receive appropriate suggestions and information that take emotions into consideration in real time, significantly improving the communication support experience. It also enables the provision of quick and appropriate answers to questions.

[0981] A "user" is an entity that utilizes the communication support system to provide voice and image data and receive support.

[0982] A "terminal" is a device that captures audio and image data in real time and transmits it to a server, and examples of this include smart glasses and smart watches.

[0983] The "server" is a central processing unit that analyzes the received audio and image data, generates appropriate suggestions, and sends them to the terminal.

[0984] The "emotion engine" is an analysis engine for recognizing the user's emotions from the user's voice data and image data.

[0985] A "generative AI model" is an artificial intelligence model that runs on a server and analyzes voice and image data to generate appropriate suggestions and answers.

[0986] "Voice data" refers to data that is a digital recording of a user's speech.

[0987] "Image data" is digital data that captures the user's face and environment with a camera.

[0988] "Analysis" is the process of extracting information from the acquired voice data and image data and understanding their meaning.

[0989] "Suggestions" are advice or information for the user that are generated by the server based on the results of the analysis.

[0990] "Question data" is voice or text data containing a specific question from the user.

[0991] An "answer" is a solution generated by the server by analyzing question data, and is information provided to the user.

[0992] "Feedback" is the act of providing suggestions or answers to a user, and can be done visually or audibly.

[0993] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this invention is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[0994] First, the user puts on a device such as smart glasses or a smart watch and turns it on. The device activates its audio and image sensors, and the user begins a conversation. The device captures the user's audio and image data in real time and sends the acquired data to a server. The device can be a common wearable device such as smart glasses or a smart watch. These devices communicate with the server via Wi-Fi or Bluetooth.

[0995] Next, the server analyzes the received voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to talk about," the emotion engine recognizes that the user is confused. Based on the emotion engine's analysis results, the server generates an appropriate suggestion, such as, "Thank you for your hard work. What would you like to talk about?" The server then sends the generated suggestion to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (audio provided through earphones).

[0996] Furthermore, if a user wants to check a specific location or object, they can capture image data of that location through the smart glasses. The device sends the captured image data to a server, which then analyzes the image data. This analysis includes cases where the user has an anxious expression, such as when searching for a conference room. The server generates navigation information such as "The conference room is further ahead on your right," and provides additional information that takes the user's emotions into consideration. The server then sends the generated information to the device, which then provides it to the user.

[0997] In addition, a real-time question-and-answer function is also provided. The user inputs a specific question by voice or text, and the device sends the question data to the server. The server analyzes the question data, and an emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression. For example, if the user is in a hurry, the server will take the urgency into account and generate a quick and detailed answer such as, "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[0998] Specific examples

[0999] Example 1: Use during a business meeting

[1000] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[1001] Example 2: Using navigation

[1002] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[1003] Prompt Sentence Examples

[1004] Below are some examples of prompts to input to the generative AI model:

[1005] "Please analyze the audio data that the user is confused about."

[1006] "Analyze image data showing a user's anxious expression."

[1007] "Generate quick answers to user questions"

[1008] As a result, users can receive appropriate communication support in real time, and difficulties in business and daily life can be alleviated.

[1009] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1010] Step 1:

[1011] The user puts on the smart glasses or smartwatch and turns on the device.

[1012] Input: A user presses the power button on the device.

[1013] Operation: The device boots up and performs self-tests and initialization.

[1014] Output: The device displays a status of ready.

[1015] Step 2:

[1016] The device activates its audio and image sensors, and the user initiates a conversation.

[1017] Input: Device readiness status and user voice input.

[1018] Action: The device activates the audio sensors and camera and begins capturing data.

[1019] Output: Audio and image data captured by the device.

[1020] Step 3:

[1021] The terminal acquires the user's voice and image data in real time.

[1022] Input: User's speech and facial expressions.

[1023] How it works: The audio sensor converts sound into digital data, and the camera captures images.

[1024] Output: Acquired audio and image data.

[1025] Step 4:

[1026] The terminal transmits the acquired voice data and image data to the server.

[1027] Input: Captured audio and image data.

[1028] How it works: Your device sends data to a server via Wi-Fi or Bluetooth.

[1029] Output: Audio and image data sent to the server.

[1030] Step 5:

[1031] The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions.

[1032] Input: Transmitted audio and image data.

[1033] How it works: The server inputs voice data into a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[1034] Output: Parsed sentiment data and appropriate suggestions.

[1035] Step 6:

[1036] The server generates appropriate suggestions based on the analysis results of the emotion engine, such as "Good work. What would you like to talk about?"

[1037] Input: Parsed emotion data and the analysis results of the generative AI model.

[1038] How it works: The server generates appropriate suggestions and converts them into text or audio format.

[1039] Output: The generated proposals.

[1040] Step 7:

[1041] The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[1042] Input: Generated proposals.

[1043] How it works: The server sends the suggestions to the device, which then provides feedback to the user, for example, on the smart glasses display or as audio through earphones.

[1044] Output: Feedback provided to the user.

[1045] Step 8:

[1046] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[1047] Input: The situation in which you are looking at the place or thing you want to check.

[1048] How it works: The user operates the smart glasses and the camera captures image data.

[1049] Output: The captured image data.

[1050] Step 9:

[1051] The terminal transmits the captured image data to the server.

[1052] Input: The captured image data.

[1053] Operation: The device compresses the image data and sends it to the server.

[1054] Output: Image data sent to the server.

[1055] Step 10:

[1056] The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expressions.

[1057] Input: The submitted image data.

[1058] How it works: The server analyzes the image data, and the emotion engine recognizes facial expressions. For example, it analyzes the facial expressions of a user searching for a meeting room.

[1059] Output: Parsed emotion data and pertinent information.

[1060] Step 11:

[1061] The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[1062] Input: Parsed emotion data and user location information.

[1063] How it works: The server generates navigation information and emotion-sensitive messages.

[1064] Output: Generated navigation information and additional information.

[1065] Step 12:

[1066] The server transmits the generated information to the terminal, which then provides it to the user.

[1067] Input: Generated navigation information and additional information.

[1068] How it works: The server sends information to the device, which then provides it to the user, for example by displaying it on smart glasses or providing audio guidance.

[1069] Output: Information provided to the user.

[1070] Step 13:

[1071] The user speaks or texts a specific question.

[1072] Input: The user's question.

[1073] How it works: A user types a question using the microphone or text input function on their smart glasses.

[1074] Output: The input question data.

[1075] Step 14:

[1076] The terminal transmits the question data to the server.

[1077] Input: The entered question data.

[1078] Operation: The device sends the query data to the server.

[1079] Output: The query data sent to the server.

[1080] Step 15:

[1081] The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression.

[1082] Input: Submitted question data and user sentiment data.

[1083] How it works: The server analyzes the question data, and the emotion engine evaluates the urgency and importance.

[1084] Output: Assessed urgency and importance.

[1085] Step 16:

[1086] The server takes into account the urgency and generates a quick and detailed response.

[1087] Input: Assessed urgency and question data.

[1088] How it works: The server uses a generative AI model to generate an answer, such as "The meaning of this word is 'Tango'."

[1089] Output: The generated answer.

[1090] Step 17:

[1091] The server generates a response and sends it to the terminal, which provides it to the user in voice or text.

[1092] Input: The generated answer.

[1093] Operation: The server sends the answer data to the terminal, which then provides it to the user via voice or text.

[1094] Output: The answer provided to the user.

[1095] (Application example 2)

[1096] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1097] Current communication support systems can make appropriate suggestions using the user's voice and image data, but they cannot generate responses based on the user's emotions, making it difficult to provide prompt and appropriate support.In addition, in certain environments such as factories, there are no systems that can properly analyze and respond to workers' difficulties and stress, making it difficult to improve productivity.

[1098] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, and means including an emotion engine for analyzing the user's emotion data and generating responses. This makes it possible to provide the user with appropriate suggestions and responses based on the emotion data, thereby achieving more effective communication support.

[1099] A "user" is an entity that uses the system to receive communication support.

[1100] "Voice data" refers to data that records the voice uttered by the user.

[1101] A "server" is a central processing unit that analyzes the data received and generates the necessary suggestions and responses.

[1102] "Emotion data" refers to data that indicates emotional information extracted from the user's voice or image.

[1103] An "emotion engine" is a system component that analyzes emotion data and recognizes the user's emotions.

[1104] A "generative AI model" is an artificial intelligence model that generates appropriate responses and suggestions based on input data.

[1105] "Suggestion" refers to the generation of useful information or advice for the user.

[1106] "Feedback" is the act of providing server-generated suggestions and responses to the user.

[1107] "Question data" is data including the content of a question provided by a user.

[1108] This invention provides a system for an emotionally responsive communication robot that supports workers in factories. Specifically, the robot acquires the worker's voice data and image data in real time, sends this data to a server for analysis, and generates appropriate suggestions and responses based on the worker's emotions. This reduces worker stress and helps improve productivity.

[1109] Hardware and software used

[1110] Hardware: A robotic device equipped with audio and image sensors, a server as the central processing unit.

[1111] Software: Speech analysis software (Google Cloud Speech-to-Text API), image analysis software (OpenCV), emotion engine (Microsoft Azure Emotion API), generative AI model (OpenAI GPT-4).

[1112] Program processing

[1113] 1. Acquisition and analysis of voice and emotion data

[1114] The worker begins working in front of the robot.

[1115] The robot activates audio and image sensors to capture the worker's voice and facial expressions in real time.

[1116] The robot transmits the acquired voice data and image data to the server.

[1117] The server analyzes the voice data using the Google Cloud Speech-to-Text API and the emotion data using the Microsoft Azure Emotion API.

[1118] Based on the emotional data, suggestions tailored to the worker's situation are generated using a generative AI model (OpenAI GPT-4).

[1119] The server sends the generated suggestions back to the robot, which then provides the suggestions to the worker via voice or display.

[1120] 2. Image Data Acquisition and Analysis

[1121] If a worker is having difficulty with a particular task, the robot will capture that situation.

[1122] The robot transmits the captured image data to a server.

[1123] The server uses OpenCV to analyze the image data, and the emotion engine recognizes the emotions from the worker's facial expressions.

[1124] Appropriate suggestions based on the worker's emotions are generated by a generative AI model, and then communicated by the robot.

[1125] Specific examples

[1126] 1. Support for workers

[1127] When a worker is performing a difficult task, the robot detects the "difficulty" from the worker's voice and facial expression.

[1128] The robot will make suggestions such as, "Thank you for your hard work. Is there anything I can help you with?"

[1129] This allows workers to receive appropriate support and carry out their work efficiently.

[1130] 2. Suggest short breaks

[1131] If a worker looks tired, the emotion engine will recognize this.

[1132] An example of a prompt from a generative AI model: "You seem to be tired. How about taking a short break?" Based on this, the model suggests, "How about taking a short break?"

[1133] This reduces worker fatigue and improves productivity.

[1134] In this way, communication support can be realized by providing appropriate suggestions and responses based on the user's emotions. This system is also very useful for supporting workers in factories and other specific environments.

[1135] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1136] Step 1:

[1137] When the user starts working in front of the robot, the robot activates the audio and image sensors. The audio sensor then captures the user's voice data, and the image sensor captures the user's facial expressions. Both data (audio and image data) are acquired in real time.

[1138] Input: User's voice and image

[1139] Output: Captured audio and image data

[1140] Step 2:

[1141] The device sends the acquired audio and image data to the server, where it is ready to be analyzed.

[1142] Input: Acquired audio and image data

[1143] Output: Audio and image data sent to the server

[1144] Step 3:

[1145] The server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then used for analysis by the emotion engine. It also uses OpenCV to extract facial expression data from the image data.

[1146] Input: Audio and image data sent to the server

[1147] Output: Analyzed voice data (text data) and facial expression data

[1148] Step 4:

[1149] The emotion engine (Microsoft Azure Emotion API) analyzes the user's emotions from the voice and facial expression data. The emotion data indicates the user's current emotional state (e.g., confusion, fatigue, tension, etc.).

[1150] Input: Analyzed voice data and facial expression data

[1151] Output: User emotion data

[1152] Step 5:

[1153] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate suggestions and responses based on the emotion data. Specific prompts are provided to the generative AI model using the emotion data as input.

[1154] Input: User emotion data

[1155] Output: The generated suggestions and responses

[1156] Step 6:

[1157] The server sends the generated suggestions and responses to the robot, which then provides the suggestions and responses to the user through voice or a display. For example, a message such as "Good work! Is there anything I can help you with?" may be displayed or spoken.

[1158] Input: Generated suggestions and responses

[1159] Output: The suggestions or responses provided to the user

[1160] Step 7:

[1161] The user receives suggestions and responses from the robot and, if necessary, further communicates with the robot. The user's reactions and additional audio and video data are processed again starting from step 1.

[1162] Input: User reactions to the suggestions and responses provided

[1163] Output: New audio and image data for the next cycle

[1164] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1165] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1166] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1167] [Third embodiment]

[1168] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1169] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1170] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1171] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1172] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1173] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1174] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1175] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1176] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1177] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1178] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1179] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1180] This invention is a communication support system that uses AI as its underlying technology, and analyzes the user's voice and image data in real time to suggest appropriate honorific expressions and words, as well as provide answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[1181] System configuration

[1182] User

[1183] This is the entity that uses the system to receive communication support. They use devices such as smart glasses or smart watches.

[1184] Terminal

[1185] A device that captures audio and image data in real time. This includes smart glasses and smart watches. The device transmits the captured data to a server.

[1186] server

[1187] It is a central processing unit that analyzes the received voice and image data. The server uses a generative AI model to analyze the data, generate appropriate suggestions and answers, and send them back to the device.

[1188] Program processing

[1189] Acquisition and analysis of audio data

[1190] 1. The user puts on the smart glasses or smartwatch and starts a conversation.

[1191] 2. The device captures the user's voice data in real time and sends the data to the server.

[1192] 3. The server analyzes the received voice data using a generative AI model. For example, if a user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates the suggestion "It's an honor to meet you."

[1193] 4. The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[1194] Image data acquisition and analysis

[1195] 1. When a user wants to see a particular place or object, they capture an image of it through the smart glasses.

[1196] 2. The device sends the captured image data to the server.

[1197] 3. The server analyzes the received image data and generates a description of the location or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on the right."

[1198] 4. The server sends the generated information to the terminal, which provides it to the user.

[1199] Real-time question answering

[1200] 1. A user asks a specific question via voice or text, for example, "What does this word mean?"

[1201] 2. The device sends the question data to the server.

[1202] 3. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'."

[1203] 4. The server generates a response and sends it to the device, which provides it to the user via voice or text.

[1204] Specific examples

[1205] Example 1: Use during a business meeting

[1206] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[1207] Example 2: Using navigation

[1208] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[1209] As described above, the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. This system is usable by a wide range of users regardless of age or level of understanding of technology, making it useful to many people.

[1210] The processing flow will be explained below.

[1211] Step 1:

[1212] The user puts on the smart glasses or smartwatch and turns on the device, which starts up the device and connects it to the network.

[1213] Step 2:

[1214] The device activates its audio and image sensors and prepares to capture data in real time.

[1215] Step 3:

[1216] When a user starts a conversation and is unsure about the meaning of a word or how to use honorific language, they speak out verbally.

[1217] Step 4:

[1218] The terminal acquires the user's voice data in real time.

[1219] Step 5:

[1220] The terminal transmits the acquired voice data to the server.

[1221] Step 6:

[1222] The server analyzes the received voice data using a generative AI model. For example, if a user says, "Oh, sorry. What was it?", the server analyzes the voice data and generates an appropriate polite expression such as, "Thank you for your hard work. What are you looking for?"

[1223] Step 7:

[1224] The server sends the analysis results to the device.

[1225] Step 8:

[1226] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[1227] Step 9:

[1228] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[1229] Step 10:

[1230] The terminal transmits the captured image data to the server.

[1231] Step 11:

[1232] The server analyzes the received image data and generates information such as "The conference room is further ahead on the right."

[1233] Step 12:

[1234] The server transmits the generated information to the terminal.

[1235] Step 13:

[1236] The device provides navigation information to the user visually (on the smart glasses display).

[1237] Step 14:

[1238] The user speaks or texts a specific question, for example, "What does this word mean?"

[1239] Step 15:

[1240] The terminal transmits the question data to the server.

[1241] Step 16:

[1242] The server analyzes the question data and generates an answer such as "The meaning of this word is 'Tango'."

[1243] Step 17:

[1244] The server sends the generated response to the terminal.

[1245] Step 18:

[1246] The device provides the user with a response by voice or text.

[1247] Step 19:

[1248] The user selects a subscription plan using a dedicated application.

[1249] Step 20:

[1250] The terminal transmits the selected plan information to the server.

[1251] Step 21:

[1252] The server enables certain features based on the plan selected by the user.

[1253] Step 22:

[1254] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[1255] This allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life.

[1256] Example 1

[1257] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1258] In modern society, communication is extremely important in business and everyday life. However, many people have difficulty using appropriate honorifics and vocabulary, and recognizing places and objects, as well as answering questions instantly, are now essential needs. Conventional communication support systems lack the ability to efficiently solve these problems in real time. In particular, more advanced technology is needed to provide appropriate content suggestions and rapid analytical feedback in a variety of situations.

[1259] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1260] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to a central processing unit, means for analyzing the voice data by the central processing unit and generating suggestions for appropriate honorific expressions and words, means for feeding back the generated suggestions to the user, and a generative AI model for analyzing the voice data. As a result, the user can receive suggestions for appropriate honorific expressions and words in real time, improving the quality of communication and reducing difficulties in business and daily life.

[1261] A "user" is an entity that uses the system to receive communication support.

[1262] "Voice data" refers to data in which the user's voice is recorded in digital form.

[1263] "Real-time" means that processing occurs immediately and on the spot with minimal delay.

[1264] "Means of acquisition" refers to devices and functions for collecting audio data and image data.

[1265] A "central processing unit" is a server or computer system that performs major processing such as data analysis.

[1266] "Means of transmission" refers to the communications means or technology used to transfer acquired data to a specific destination.

[1267] "Means for analyzing" refers to software or algorithms used to analyze received data.

[1268] "Means for generating suggestions" refers to a function that generates appropriate information and advice to be provided to the user based on the analyzed data.

[1269] "Feedback means" refers to an output device or method for notifying the user of the generated suggestions.

[1270] "Image data" is data that records a user's visual information in digital form.

[1271] "Location description" refers to content that provides information or directions to a specific location.

[1272] "Name of object" refers to the name of the object contained in the image data.

[1273] "Question data" refers to data that indicates the content of a question that a user asks the system.

[1274] "Means for generating answers" refers to algorithms or software for generating optimal answers based on question data.

[1275] A "generative AI model" is a model or system that uses artificial intelligence to analyze data and generate suggestions or answers.

[1276] This invention is a communication support system that uses a generative AI model as its underlying technology. It analyzes the user's voice and image data in real time, and suggests appropriate honorific expressions and words, as well as providing answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[1277] Acquisition and analysis of audio data

[1278] A user puts on smart glasses or a smartwatch and starts a conversation. The device captures the user's voice data in real time and sends it to a server. The server analyzes the received voice data using a generative AI model (e.g., OpenAI's GPT-4). For example, if the user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates a suggestion: "It's an honor to meet you." The server sends the suggestion to the device, which then provides visual or auditory feedback to the user.

[1279] Image data acquisition and analysis

[1280] When a user wants to check a specific place or object, they capture an image of it through the smart glasses. The device sends the captured image data to a server. The server analyzes the received image data and generates a description of the place or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on your right." The server sends the generated information to the device, which then provides it to the user.

[1281] Real-time question answering

[1282] The user asks a specific question by voice or text. For example, "What does this word mean?", the device sends the question data to the server. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[1283] Specific examples

[1284] Example 1: Use during a business meeting

[1285] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[1286] Example 2: Using navigation

[1287] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[1288] Prompt Sentence Examples

[1289] Use the following prompt:

[1290] User Input: "Nice to meet you."

[1291] AI Model: "It's a pleasure to meet you."

[1292] User Input: "What does this word mean?"

[1293] AI Model: "This word means 'Tango'"

[1294] In this way, the system of the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. Furthermore, by using a generative AI model, it is possible to respond quickly and accurately to user needs, making the system useful for many people, regardless of age or level of technological literacy.

[1295] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1296] Acquisition and analysis of audio data

[1297] Step 1: The user puts on the smart glasses or smartwatch and starts a conversation.

[1298] Input: User's voice

[1299] Action: The user speaks.

[1300] Output: Audio signal

[1301] Step 2: The terminal acquires the user's voice data in real time.

[1302] Input: Audio signal

[1303] How it works: The device's microphone picks up audio signals and converts them into digital audio data.

[1304] Output: Digital audio data

[1305] Step 3: The device sends the acquired voice data to the server.

[1306] Input: Digital audio data

[1307] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[1308] Output: Compressed and encrypted audio data

[1309] Step 4: The server analyzes the received voice data using the generative AI model.

[1310] Input: Compressed and encrypted audio data

[1311] How it works: The server decodes the audio data and uses a generative AI model (e.g., GPT-4) to convert the audio to text and perform analysis.

[1312] Output: Analysis results (suggestions for appropriate honorifics and words)

[1313] Step 5: The server sends the generated proposal to the device.

[1314] Input: Analysis results (suggestions for appropriate honorifics and words)

[1315] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[1316] Output: Compressed and encrypted proposal data

[1317] Step 6: The device provides visual or auditory feedback to the user.

[1318] Input: Compressed and encrypted proposal data

[1319] How it works: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[1320] Output: User feedback (visual / auditory information)

[1321] Image data acquisition and analysis

[1322] Step 1: When a user wants to see a specific place or object, they capture an image of it through their smart glasses.

[1323] Input: Visual information

[1324] Action: The user operates the smartglasses camera to capture an image.

[1325] Output: Image data

[1326] Step 2: The device sends the captured image data to the server.

[1327] Input: Image data

[1328] Operation: The terminal's communication module compresses and encrypts the image data and sends it to the server.

[1329] Output: Compressed and encrypted image data

[1330] Step 3: The server analyzes the received image data.

[1331] Input: Compressed and encrypted image data

[1332] How it works: The server decodes the data and uses image recognition software to analyze the image (e.g., recognize location descriptions and object names).

[1333] Output: Analysis results (location descriptions and names of things)

[1334] Step 4: The server sends the generated information to the terminal.

[1335] Input: Analysis results (location description and object names)

[1336] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[1337] Output: Compressed and encrypted information data

[1338] Step 5: The terminal provides the user with information visually or audibly.

[1339] Input: Compressed and encrypted information data

[1340] How it works: The device decodes the data and displays it on the screen or speaks the information over the speaker.

[1341] Output: Provide information to the user (visual / auditory information)

[1342] Real-time question answering

[1343] Step 1: The user asks a specific question via voice or text.

[1344] Input: Voice or text question

[1345] How it works: If the user asks a question by voice, the device's microphone picks up the sound, and if the user asks a question by text, the device's text input function is used.

[1346] Output: Digital audio data or text data

[1347] Step 2: The device sends the query data to the server.

[1348] Input: Digital audio or text data

[1349] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[1350] Output: Compressed and encrypted query data

[1351] Step 3: The server analyzes the question data and generates the best answer.

[1352] Input: Compressed and encrypted query data

[1353] How it works: The server decodes the data, uses a generative AI model to analyze the intent of the question, and generates an answer.

[1354] Output: Analysis results (best answer)

[1355] Step 4: The server sends the generated response to the terminal.

[1356] Input: Analysis results (best answer)

[1357] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[1358] Output: Compressed and encrypted answer data

[1359] Step 5: The device provides the user with a voice or text message.

[1360] Input: Compressed and encrypted answer data

[1361] What it does: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[1362] Output: Provides answers to users (visual / auditory information)

[1363] As described above, the system achieves efficient communication support for users through specific operations and data input / output at each step.

[1364] (Application example 1)

[1365] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1366] Conventional communication support systems allow users to receive suggestions for appropriate honorific expressions and words, but they have difficulty answering customer questions in real time or providing in-store navigation information when dealing with customers in brick-and-mortar stores. They also lack the ability to provide personalized services based on customer facial recognition or past purchase history. The present invention provides a new communication support system that streamlines customer service in brick-and-mortar stores and improves customer satisfaction.

[1367] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1368] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, means for analyzing question data acquired from the user via the smart wearable device and generating an optimal answer using a generative AI model, means for visually or audibly feeding back the generated answer, and means for displaying in-store navigation information based on the generated information, thereby enabling immediate responses to customer questions in physical stores, in-store navigation, and personalized services.

[1369] "Voice data" is information about sounds generated by the user's vocalizations.

[1370] A "server" is a central processing unit that processes and provides various data via a network.

[1371] "Analysis" is the act of interpreting and understanding the content of acquired data to generate appropriate information.

[1372] A "generative AI model" is a trained algorithm that uses artificial intelligence technology to analyze data and generate appropriate suggestions and answers.

[1373] Honorific language is the use of words to show respect to the other person, and is part of language usage.

[1374] "Word suggestion" is the act of selecting appropriate words and providing them to the user.

[1375] A "smart wearable device" is a portable electronic device that is worn by a user.

[1376] "Question data" is data of statements or text that a user uses to ask for specific information.

[1377] "Visual feedback" is a method of providing information to users in a form that they can see.

[1378] "Auditory feedback" is a method of providing information to a user in a form that can be heard by ear.

[1379] "Navigation information" is information that provides directions and location information to a specific location.

[1380] "Feedback" is the act of returning information or results from a system to a user.

[1381] The present invention is a communication support system for improving the efficiency of customer service in brick-and-mortar stores and enhancing customer satisfaction. This system is composed of three main elements: a user, a terminal (a smart wearable device), and a server. Specific embodiments for implementing the present invention are described below.

[1382] Acquisition and analysis of audio data

[1383] The user wears a device such as smart glasses or a smartwatch and responds to customers. The device captures the user's voice data in real time and sends that data to a server. The server then analyzes the voice data using a generative AI model and generates appropriate honorific and word suggestions. The generated suggestions are then fed back to the user visually or audibly via the device.

[1384] Image data acquisition and analysis

[1385] The user acquires image data of the interior of a store and the products through the device. The device then sends the captured image data to the server. The server analyzes the image data and generates location descriptions and names of objects. For example, it can also provide navigation information such as "The conference room is further ahead on the right."

[1386] Real-time question answering

[1387] When a user receives a question from a customer, the device sends the question data as voice or text to the server. The server analyzes the question data using a generative AI model and generates the optimal answer. This answer is also fed back to the user via the device.

[1388] Specific examples from physical stores

[1389] Example 1: Customer Service

[1390] When a customer asks, "Tell me about this new product," the user's device sends the question to the server, which uses a generative AI model to generate an answer such as, "This new product uses the latest AI technology and has a user-friendly interface," and the device visually displays the answer to the user.

[1391] Example 2: In-store navigation

[1392] When a customer asks, "Where is the conference room?", the user captures an image of the surroundings through the smart glasses and sends it to the server. The server analyzes the image data and generates navigation information such as "The conference room is on the second floor, on the right," which is then visually presented to the user on the device.

[1393] Hardware and software used

[1394] Hardware:

[1395] Smart Glasses

[1396] Smartwatch

[1397] microphone

[1398] software:

[1399] Python

[1400] SpeechRecognition Library

[1401] PIL (Python Imaging Library)

[1402] Requests library

[1403] OpenAI API

[1404] Prompt Sentence Examples

[1405] Here are some examples of prompts you can give to your generative AI model:

[1406] Customer Question: What are the features of this product?

[1407] Generate the appropriate answer.

[1408]

[1409] Please describe the image. The image shows a conference room.

[1410] The present invention allows users in physical stores to smoothly handle customer inquiries, provide instant responses to questions, and provide in-store navigation, which is expected to improve the quality of service and increase customer satisfaction.

[1411] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1412] Step 1:

[1413] The user puts on the smart glasses or smartwatch and begins interacting with the customer.

[1414] Input: User's voice or image data

[1415] Output: Captured audio or image data

[1416] Specific operation: The device's microphone captures audio data in real time, and the camera captures image data.

[1417] Step 2:

[1418] The terminal transmits the acquired audio and image data to the server.

[1419] Input: Audio or image data captured by the device

[1420] Output: Audio or image data transferred to the server

[1421] Specific operation: The terminal transmits the collected audio or image data to the server via the network.

[1422] Step 3:

[1423] The server analyzes the audio data and generates appropriate honorifics and word suggestions.

[1424] Input: Audio data received by the server

[1425] Output: Generated honorifics and word suggestions

[1426] Specific operation: The server processes the speech data using a generative AI model to generate appropriate honorifics and word suggestions.

[1427] Step 4:

[1428] The server analyzes the image data and generates location descriptions and names of objects.

[1429] Input: Image data received by the server

[1430] Output: Generated place descriptions and object names

[1431] How it works: The server processes image data using a generative AI model to identify location descriptions and object names.

[1432] Step 5:

[1433] The server sends generated suggestions and information to the terminal, which then provides visual or auditory feedback to the user.

[1434] Input: Generated suggestions and information from the server

[1435] Output: Feedback information provided on the device by display or audio

[1436] Specific operation: The server sends the generated honorific suggestions and navigation information to the device, which then displays them on the screen or notifies the user by voice.

[1437] Step 6:

[1438] It analyzes question data obtained from users via smart wearable devices and generates optimal answers using a generative AI model.

[1439] Input: Question data from the user

[1440] Output: The generated answer

[1441] How it works: The device receives a question via voice or text and sends it to the server, which then uses a generative AI model to analyze the question data and generate an answer.

[1442] Step 7:

[1443] The server generates a response and sends it to the terminal, which then provides visual or auditory feedback to the user.

[1444] Input: Generated answer from the server

[1445] Output: Answer information displayed on the device or provided as audio

[1446] Specific operation: The answer generated by the server is sent to the terminal and presented in a format that is easy for the user to see or hear.

[1447] Step 8:

[1448] The server displays in-store navigation information based on the generated information.

[1449] Input: Location information and image data from the user

[1450] Output: Generated navigation information

[1451] Specific operation: The server generates in-store navigation information based on location information and image data and sends it to the device, which then visually guides the user along the way.

[1452] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1453] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this embodiment is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[1454] System configuration

[1455] User

[1456] This is the subject who uses the system to receive communication support. They wear devices such as smart glasses or smart watches.

[1457] Terminal

[1458] A device that captures audio and image data in real time. In this case, it refers to smart glasses and smart watches. The device is responsible for sending the captured data to a server.

[1459] server

[1460] This is a central processing unit that analyzes the received voice and image data. It uses generative AI models to analyze the data and generate appropriate suggestions, which are then sent back to the device.

[1461] Emotion Engine

[1462] This engine recognizes the user's emotions from their voice and image data. Based on the emotion data, the server generates appropriate suggestions.

[1463] Program processing

[1464] Acquisition and analysis of voice and emotion data

[1465] 1. The user puts on the smart glasses or smartwatch and turns on the device.

[1466] 2. The device activates its audio and image sensors, and the user initiates a conversation.

[1467] 3. The terminal acquires the user's voice and image data in real time.

[1468] 4. The terminal sends the acquired voice and image data to the server.

[1469] 5. The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[1470] 6. The server generates appropriate suggestions based on the emotion engine analysis, such as "Great work. What would you like to talk about?"

[1471] 7. The server sends the generated suggestions to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (sound presented through earphones).

[1472] Image data acquisition and analysis

[1473] 1. When a user wants to check a specific place or object, they capture its image data through smart glasses.

[1474] 2. The device sends the captured image data to the server.

[1475] 3. The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expression. For example, if a user looks anxious while searching for a conference room.

[1476] 4. The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[1477] 5. The server sends the generated information to the terminal, which provides it to the user.

[1478] Real-time question answering

[1479] 1. The user speaks or writes a specific question, for example, "What does this word mean?"

[1480] 2. The device sends the question data to the server.

[1481] 3. The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expressions. For example, if the user is in a hurry.

[1482] 4. The server generates a quick and detailed answer based on the urgency of the question, such as "The meaning of this word is 'Tango'."

[1483] 5. The server sends the generated answer to the device, which provides it to the user via voice or text.

[1484] Specific examples

[1485] Example 1: Use during a business meeting

[1486] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[1487] Example 2: Using navigation

[1488] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[1489] As described above, the present invention allows users to receive appropriate communication support in real time, thereby reducing difficulties in business and daily life. This system is available to a wide range of users, regardless of age or level of understanding of technology, and is useful to many people.

[1490] The processing flow will be explained below.

[1491] Step 1:

[1492] The user puts on the smart glasses or smart watch and turns on the device. The device starts up and connects to the network.

[1493] Step 2:

[1494] The device activates its audio and image sensors and prepares to capture audio and images in real time.

[1495] Step 3:

[1496] When a user starts a conversation and has difficulty choosing words or using honorific language, they speak out verbally.

[1497] Step 4:

[1498] The terminal acquires the user's voice data in real time.

[1499] Step 5:

[1500] The terminal transmits the acquired voice data to the server.

[1501] Step 6:

[1502] The server analyzes the received voice data using a generative AI model.

[1503] Step 7:

[1504] The emotion engine recognizes the user's emotion based on the voice data. For example, if the user says, "I'm sorry, I'm in trouble," the emotion engine recognizes the user's confusion.

[1505] Step 8:

[1506] The server generates appropriate honorifics and word suggestions based on the analysis results of the emotion engine. For example, it generates a suggestion such as, "You seem to be in trouble. What can I explain to you?"

[1507] Step 9:

[1508] The server sends the generated proposal to the terminal.

[1509] Step 10:

[1510] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[1511] Step 11:

[1512] When a user wants to view a particular object or location, they capture image data of it through the smart glasses.

[1513] Step 12:

[1514] The terminal transmits the captured image data to the server.

[1515] Step 13:

[1516] The server analyzes the image data.

[1517] Step 14:

[1518] The emotion engine recognizes the user's emotional state from their facial expressions, for example, when the user has an anxious expression.

[1519] Step 15:

[1520] The server generates information including a description of the location and the name of the object based on the image analysis results and the recognition results of the emotion engine. For example, it could generate information such as "This is the conference room. It's up ahead on your right," providing additional information to ease the user's anxiety.

[1521] Step 16:

[1522] The server transmits the generated information to the terminal.

[1523] Step 17:

[1524] The device provides navigation information to the user visually (on the smart glasses display).

[1525] Step 18:

[1526] The user speaks or texts a specific question.

[1527] Step 19:

[1528] The terminal transmits the question data to the server.

[1529] Step 20:

[1530] The server parses the question data.

[1531] Step 21:

[1532] The emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression when asking the question. For example, if it determines that the question is urgent,

[1533] Step 22:

[1534] The server generates a quick and detailed answer based on the evaluation of the emotion engine, for example, "The meaning of this word is 'Tango'."

[1535] Step 23:

[1536] The server generates a response and sends it to the terminal.

[1537] Step 24:

[1538] The device provides the user with a response by voice or text.

[1539] Step 25:

[1540] The user selects a subscription plan using a dedicated application.

[1541] Step 26:

[1542] The terminal transmits the selected plan information to the server.

[1543] Step 27:

[1544] The server enables certain features based on the plan selected by the user.

[1545] Step 28:

[1546] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[1547] Through this process, the system enables users to receive appropriate communication assistance in real time, easing difficulties in business and daily life.

[1548] Example 2

[1549] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1550] Conventional communication support systems have had difficulty in providing real-time support by fully utilizing the user's voice data and image data. In particular, they lacked the ability to provide appropriate suggestions and information that took the user's emotions into consideration, making it difficult to improve the user experience. They also lacked the ability to provide prompt and appropriate answers to user questions. It is necessary to provide a system that can solve these problems and enable users to receive better support.

[1551] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1552] In this invention, the server includes means for acquiring and processing a user's voice data and image data in real time, means for transmitting the voice data and image data to the server, means for analyzing the voice data and image data by the server and recognizing emotions, means for analyzing the voice data using a generative AI model to generate appropriate suggestions, and means for feeding back the generated suggestions to the user. This allows the user to receive appropriate suggestions and information that take emotions into consideration in real time, significantly improving the communication support experience. It also enables the provision of quick and appropriate answers to questions.

[1553] A "user" is an entity that utilizes the communication support system to provide voice and image data and receive support.

[1554] A "terminal" is a device that captures audio and image data in real time and transmits it to a server, and examples of this include smart glasses and smart watches.

[1555] The "server" is a central processing unit that analyzes the received audio and image data, generates appropriate suggestions, and sends them to the terminal.

[1556] The "emotion engine" is an analysis engine for recognizing the user's emotions from the user's voice data and image data.

[1557] A "generative AI model" is an artificial intelligence model that runs on a server and analyzes voice and image data to generate appropriate suggestions and answers.

[1558] "Voice data" refers to data that is a digital recording of a user's speech.

[1559] "Image data" is digital data that captures the user's face and environment with a camera.

[1560] "Analysis" is the process of extracting information from the acquired voice data and image data and understanding their meaning.

[1561] "Suggestions" are advice or information for the user that are generated by the server based on the results of the analysis.

[1562] "Question data" is voice or text data containing a specific question from the user.

[1563] An "answer" is a solution generated by the server by analyzing question data, and is information provided to the user.

[1564] "Feedback" is the act of providing suggestions or answers to a user, and can be done visually or audibly.

[1565] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this invention is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[1566] First, the user puts on a device such as smart glasses or a smart watch and turns it on. The device activates its audio and image sensors, and the user begins a conversation. The device captures the user's audio and image data in real time and sends the acquired data to a server. The device can be a common wearable device such as smart glasses or a smart watch. These devices communicate with the server via Wi-Fi or Bluetooth.

[1567] Next, the server analyzes the received voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to talk about," the emotion engine recognizes that the user is confused. Based on the emotion engine's analysis results, the server generates an appropriate suggestion, such as, "Thank you for your hard work. What would you like to talk about?" The server then sends the generated suggestion to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (audio provided through earphones).

[1568] Furthermore, if a user wants to check a specific location or object, they can capture image data of that location through the smart glasses. The device sends the captured image data to a server, which then analyzes the image data. This analysis includes cases where the user has an anxious expression, such as when searching for a conference room. The server generates navigation information such as "The conference room is further ahead on your right," and provides additional information that takes the user's emotions into consideration. The server then sends the generated information to the device, which then provides it to the user.

[1569] In addition, a real-time question-and-answer function is also provided. The user inputs a specific question by voice or text, and the device sends the question data to the server. The server analyzes the question data, and an emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression. For example, if the user is in a hurry, the server will take the urgency into account and generate a quick and detailed answer such as, "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[1570] Specific examples

[1571] Example 1: Use during a business meeting

[1572] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[1573] Example 2: Using navigation

[1574] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[1575] Prompt Sentence Examples

[1576] Below are some examples of prompts to input to the generative AI model:

[1577] "Please analyze the audio data that the user is confused about."

[1578] "Analyze image data showing a user's anxious expression."

[1579] "Generate quick answers to user questions"

[1580] As a result, users can receive appropriate communication support in real time, and difficulties in business and daily life can be alleviated.

[1581] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1582] Step 1:

[1583] The user puts on the smart glasses or smartwatch and turns on the device.

[1584] Input: A user presses the power button on the device.

[1585] Operation: The device boots up and performs self-tests and initialization.

[1586] Output: The device displays a status of ready.

[1587] Step 2:

[1588] The device activates its audio and image sensors, and the user initiates a conversation.

[1589] Input: Device readiness status and user voice input.

[1590] Action: The device activates the audio sensors and camera and begins capturing data.

[1591] Output: Audio and image data captured by the device.

[1592] Step 3:

[1593] The terminal acquires the user's voice and image data in real time.

[1594] Input: User's speech and facial expressions.

[1595] How it works: The audio sensor converts sound into digital data, and the camera captures images.

[1596] Output: Acquired audio and image data.

[1597] Step 4:

[1598] The terminal transmits the acquired voice data and image data to the server.

[1599] Input: Captured audio and image data.

[1600] How it works: Your device sends data to a server via Wi-Fi or Bluetooth.

[1601] Output: Audio and image data sent to the server.

[1602] Step 5:

[1603] The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions.

[1604] Input: Transmitted audio and image data.

[1605] How it works: The server inputs voice data into a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[1606] Output: Parsed sentiment data and appropriate suggestions.

[1607] Step 6:

[1608] The server generates appropriate suggestions based on the analysis results of the emotion engine, such as "Good work. What would you like to talk about?"

[1609] Input: Parsed emotion data and the analysis results of the generative AI model.

[1610] How it works: The server generates appropriate suggestions and converts them into text or audio format.

[1611] Output: The generated proposals.

[1612] Step 7:

[1613] The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[1614] Input: Generated proposals.

[1615] How it works: The server sends the suggestions to the device, which then provides feedback to the user, for example, on the smart glasses display or as audio through earphones.

[1616] Output: Feedback provided to the user.

[1617] Step 8:

[1618] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[1619] Input: The situation in which you are looking at the place or thing you want to check.

[1620] How it works: The user operates the smart glasses and the camera captures image data.

[1621] Output: The captured image data.

[1622] Step 9:

[1623] The terminal transmits the captured image data to the server.

[1624] Input: The captured image data.

[1625] Operation: The device compresses the image data and sends it to the server.

[1626] Output: Image data sent to the server.

[1627] Step 10:

[1628] The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expressions.

[1629] Input: The submitted image data.

[1630] How it works: The server analyzes the image data, and the emotion engine recognizes facial expressions. For example, it analyzes the facial expressions of a user searching for a meeting room.

[1631] Output: Parsed emotion data and pertinent information.

[1632] Step 11:

[1633] The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[1634] Input: Parsed emotion data and user location information.

[1635] How it works: The server generates navigation information and emotion-sensitive messages.

[1636] Output: Generated navigation information and additional information.

[1637] Step 12:

[1638] The server transmits the generated information to the terminal, which then provides it to the user.

[1639] Input: Generated navigation information and additional information.

[1640] How it works: The server sends information to the device, which then provides it to the user, for example by displaying it on smart glasses or providing audio guidance.

[1641] Output: Information provided to the user.

[1642] Step 13:

[1643] The user speaks or texts a specific question.

[1644] Input: The user's question.

[1645] How it works: A user types a question using the microphone or text input function on their smart glasses.

[1646] Output: The input question data.

[1647] Step 14:

[1648] The terminal transmits the question data to the server.

[1649] Input: The entered question data.

[1650] Operation: The device sends the query data to the server.

[1651] Output: The query data sent to the server.

[1652] Step 15:

[1653] The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression.

[1654] Input: Submitted question data and user sentiment data.

[1655] How it works: The server analyzes the question data, and the emotion engine evaluates the urgency and importance.

[1656] Output: Assessed urgency and importance.

[1657] Step 16:

[1658] The server takes into account the urgency and generates a quick and detailed response.

[1659] Input: Assessed urgency and question data.

[1660] How it works: The server uses a generative AI model to generate an answer, such as "The meaning of this word is 'Tango'."

[1661] Output: The generated answer.

[1662] Step 17:

[1663] The server generates a response and sends it to the terminal, which provides it to the user in voice or text.

[1664] Input: The generated answer.

[1665] Operation: The server sends the answer data to the terminal, which then provides it to the user via voice or text.

[1666] Output: The answer provided to the user.

[1667] (Application example 2)

[1668] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1669] Current communication support systems can make appropriate suggestions using the user's voice and image data, but they cannot generate responses based on the user's emotions, making it difficult to provide prompt and appropriate support.In addition, in certain environments such as factories, there are no systems that can properly analyze and respond to workers' difficulties and stress, making it difficult to improve productivity.

[1670] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, and means including an emotion engine for analyzing the user's emotion data and generating responses. This makes it possible to provide the user with appropriate suggestions and responses based on the emotion data, thereby achieving more effective communication support.

[1671] A "user" is an entity that uses the system to receive communication support.

[1672] "Voice data" refers to data that records the voice uttered by the user.

[1673] A "server" is a central processing unit that analyzes the data received and generates the necessary suggestions and responses.

[1674] "Emotion data" refers to data that indicates emotional information extracted from the user's voice or image.

[1675] An "emotion engine" is a system component that analyzes emotion data and recognizes the user's emotions.

[1676] A "generative AI model" is an artificial intelligence model that generates appropriate responses and suggestions based on input data.

[1677] "Suggestion" refers to the generation of useful information or advice for the user.

[1678] "Feedback" is the act of providing server-generated suggestions and responses to the user.

[1679] "Question data" is data including the content of a question provided by a user.

[1680] This invention provides a system for an emotionally responsive communication robot that supports workers in factories. Specifically, the robot acquires the worker's voice data and image data in real time, sends this data to a server for analysis, and generates appropriate suggestions and responses based on the worker's emotions. This reduces worker stress and helps improve productivity.

[1681] Hardware and software used

[1682] Hardware: A robotic device equipped with audio and image sensors, a server as the central processing unit.

[1683] Software: Speech analysis software (Google Cloud Speech-to-Text API), image analysis software (OpenCV), emotion engine (Microsoft Azure Emotion API), generative AI model (OpenAI GPT-4).

[1684] Program processing

[1685] 1. Acquisition and analysis of voice and emotion data

[1686] The worker begins working in front of the robot.

[1687] The robot activates audio and image sensors to capture the worker's voice and facial expressions in real time.

[1688] The robot transmits the acquired voice data and image data to the server.

[1689] The server analyzes the voice data using the Google Cloud Speech-to-Text API and the emotion data using the Microsoft Azure Emotion API.

[1690] Based on the emotional data, suggestions tailored to the worker's situation are generated using a generative AI model (OpenAI GPT-4).

[1691] The server sends the generated suggestions back to the robot, which then provides the suggestions to the worker via voice or display.

[1692] 2. Image Data Acquisition and Analysis

[1693] If a worker is having difficulty with a particular task, the robot will capture that situation.

[1694] The robot transmits the captured image data to a server.

[1695] The server uses OpenCV to analyze the image data, and the emotion engine recognizes the emotions from the worker's facial expressions.

[1696] Appropriate suggestions based on the worker's emotions are generated by a generative AI model, and then communicated by the robot.

[1697] Specific examples

[1698] 1. Support for workers

[1699] When a worker is performing a difficult task, the robot detects the "difficulty" from the worker's voice and facial expression.

[1700] The robot will make suggestions such as, "Thank you for your hard work. Is there anything I can help you with?"

[1701] This allows workers to receive appropriate support and carry out their work efficiently.

[1702] 2. Suggest short breaks

[1703] If a worker looks tired, the emotion engine will recognize this.

[1704] An example of a prompt from a generative AI model: "You seem to be tired. How about taking a short break?" Based on this, the model suggests, "How about taking a short break?"

[1705] This reduces worker fatigue and improves productivity.

[1706] In this way, communication support can be realized by providing appropriate suggestions and responses based on the user's emotions. This system is also very useful for supporting workers in factories and other specific environments.

[1707] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1708] Step 1:

[1709] When the user starts working in front of the robot, the robot activates the audio and image sensors. The audio sensor then captures the user's voice data, and the image sensor captures the user's facial expressions. Both data (audio and image data) are acquired in real time.

[1710] Input: User's voice and image

[1711] Output: Captured audio and image data

[1712] Step 2:

[1713] The device sends the acquired audio and image data to the server, where it is ready to be analyzed.

[1714] Input: Acquired audio and image data

[1715] Output: Audio and image data sent to the server

[1716] Step 3:

[1717] The server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then used for analysis by the emotion engine. It also uses OpenCV to extract facial expression data from the image data.

[1718] Input: Audio and image data sent to the server

[1719] Output: Analyzed voice data (text data) and facial expression data

[1720] Step 4:

[1721] The emotion engine (Microsoft Azure Emotion API) analyzes the user's emotions from the voice and facial expression data. The emotion data indicates the user's current emotional state (e.g., confusion, fatigue, tension, etc.).

[1722] Input: Analyzed voice data and facial expression data

[1723] Output: User emotion data

[1724] Step 5:

[1725] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate suggestions and responses based on the emotion data. Specific prompts are provided to the generative AI model using the emotion data as input.

[1726] Input: User emotion data

[1727] Output: The generated suggestions and responses

[1728] Step 6:

[1729] The server sends the generated suggestions and responses to the robot, which then provides the suggestions and responses to the user through voice or a display. For example, a message such as "Good work! Is there anything I can help you with?" may be displayed or spoken.

[1730] Input: Generated suggestions and responses

[1731] Output: The suggestions or responses provided to the user

[1732] Step 7:

[1733] The user receives suggestions and responses from the robot and, if necessary, further communicates with the robot. The user's reactions and additional audio and video data are processed again starting from step 1.

[1734] Input: User reactions to the suggestions and responses provided

[1735] Output: New audio and image data for the next cycle

[1736] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1737] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1738] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1739] [Fourth embodiment]

[1740] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1741] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1742] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1743] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1744] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1745] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1746] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1747] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1748] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1749] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1750] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1751] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1752] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1753] This invention is a communication support system that uses AI as its underlying technology, and analyzes the user's voice and image data in real time to suggest appropriate honorific expressions and words, as well as provide answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[1754] System configuration

[1755] User

[1756] This is the entity that uses the system to receive communication support. They use devices such as smart glasses or smart watches.

[1757] Terminal

[1758] A device that captures audio and image data in real time. This includes smart glasses and smart watches. The device transmits the captured data to a server.

[1759] server

[1760] It is a central processing unit that analyzes the received voice and image data. The server uses a generative AI model to analyze the data, generate appropriate suggestions and answers, and send them back to the device.

[1761] Program processing

[1762] Acquisition and analysis of audio data

[1763] 1. The user puts on the smart glasses or smartwatch and starts a conversation.

[1764] 2. The device captures the user's voice data in real time and sends the data to the server.

[1765] 3. The server analyzes the received voice data using a generative AI model. For example, if a user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates the suggestion "It's an honor to meet you."

[1766] 4. The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[1767] Image data acquisition and analysis

[1768] 1. When a user wants to see a particular place or object, they capture an image of it through the smart glasses.

[1769] 2. The device sends the captured image data to the server.

[1770] 3. The server analyzes the received image data and generates a description of the location or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on the right."

[1771] 4. The server sends the generated information to the terminal, which provides it to the user.

[1772] Real-time question answering

[1773] 1. A user asks a specific question via voice or text, for example, "What does this word mean?"

[1774] 2. The device sends the question data to the server.

[1775] 3. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'."

[1776] 4. The server generates a response and sends it to the device, which provides it to the user via voice or text.

[1777] Specific examples

[1778] Example 1: Use during a business meeting

[1779] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[1780] Example 2: Using navigation

[1781] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[1782] As described above, the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. This system is usable by a wide range of users regardless of age or level of understanding of technology, making it useful to many people.

[1783] The processing flow will be explained below.

[1784] Step 1:

[1785] The user puts on the smart glasses or smartwatch and turns on the device, which starts up the device and connects it to the network.

[1786] Step 2:

[1787] The device activates its audio and image sensors and prepares to capture data in real time.

[1788] Step 3:

[1789] When a user starts a conversation and is unsure about the meaning of a word or how to use honorific language, they speak out verbally.

[1790] Step 4:

[1791] The terminal acquires the user's voice data in real time.

[1792] Step 5:

[1793] The terminal transmits the acquired voice data to the server.

[1794] Step 6:

[1795] The server analyzes the received voice data using a generative AI model. For example, if a user says, "Oh, sorry. What was it?", the server analyzes the voice data and generates an appropriate polite expression such as, "Thank you for your hard work. What are you looking for?"

[1796] Step 7:

[1797] The server sends the analysis results to the device.

[1798] Step 8:

[1799] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[1800] Step 9:

[1801] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[1802] Step 10:

[1803] The terminal transmits the captured image data to the server.

[1804] Step 11:

[1805] The server analyzes the received image data and generates information such as "The conference room is further ahead on the right."

[1806] Step 12:

[1807] The server transmits the generated information to the terminal.

[1808] Step 13:

[1809] The device provides navigation information to the user visually (on the smart glasses display).

[1810] Step 14:

[1811] The user speaks or texts a specific question, for example, "What does this word mean?"

[1812] Step 15:

[1813] The terminal transmits the question data to the server.

[1814] Step 16:

[1815] The server analyzes the question data and generates an answer such as "The meaning of this word is 'Tango'."

[1816] Step 17:

[1817] The server sends the generated response to the terminal.

[1818] Step 18:

[1819] The device provides the user with a response by voice or text.

[1820] Step 19:

[1821] The user selects a subscription plan using a dedicated application.

[1822] Step 20:

[1823] The terminal transmits the selected plan information to the server.

[1824] Step 21:

[1825] The server enables certain features based on the plan selected by the user.

[1826] Step 22:

[1827] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[1828] This allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life.

[1829] Example 1

[1830] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1831] In modern society, communication is extremely important in business and everyday life. However, many people have difficulty using appropriate honorifics and vocabulary, and recognizing places and objects, as well as answering questions instantly, are now essential needs. Conventional communication support systems lack the ability to efficiently solve these problems in real time. In particular, more advanced technology is needed to provide appropriate content suggestions and rapid analytical feedback in a variety of situations.

[1832] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1833] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to a central processing unit, means for analyzing the voice data by the central processing unit and generating suggestions for appropriate honorific expressions and words, means for feeding back the generated suggestions to the user, and a generative AI model for analyzing the voice data. As a result, the user can receive suggestions for appropriate honorific expressions and words in real time, improving the quality of communication and reducing difficulties in business and daily life.

[1834] A "user" is an entity that uses the system to receive communication support.

[1835] "Voice data" refers to data in which the user's voice is recorded in digital form.

[1836] "Real-time" means that processing occurs immediately and on the spot with minimal delay.

[1837] "Means of acquisition" refers to devices and functions for collecting audio data and image data.

[1838] A "central processing unit" is a server or computer system that performs major processing such as data analysis.

[1839] "Means of transmission" refers to the communications means or technology used to transfer acquired data to a specific destination.

[1840] "Means for analyzing" refers to software or algorithms used to analyze received data.

[1841] "Means for generating suggestions" refers to a function that generates appropriate information and advice to be provided to the user based on the analyzed data.

[1842] "Feedback means" refers to an output device or method for notifying the user of the generated suggestions.

[1843] "Image data" is data that records a user's visual information in digital form.

[1844] "Location description" refers to content that provides information or directions to a specific location.

[1845] "Name of object" refers to the name of the object contained in the image data.

[1846] "Question data" refers to data that indicates the content of a question that a user asks the system.

[1847] "Means for generating answers" refers to algorithms or software for generating optimal answers based on question data.

[1848] A "generative AI model" is a model or system that uses artificial intelligence to analyze data and generate suggestions or answers.

[1849] This invention is a communication support system that uses a generative AI model as its underlying technology. It analyzes the user's voice and image data in real time, and suggests appropriate honorific expressions and words, as well as providing answers to questions and navigation information. The system of this embodiment is composed of three main elements: the user, the terminal, and the server.

[1850] Acquisition and analysis of audio data

[1851] A user puts on smart glasses or a smartwatch and starts a conversation. The device captures the user's voice data in real time and sends it to a server. The server analyzes the received voice data using a generative AI model (e.g., OpenAI's GPT-4). For example, if the user wants to say "It's nice to meet you" but doesn't know the appropriate honorific, the server generates a suggestion: "It's an honor to meet you." The server sends the suggestion to the device, which then provides visual or auditory feedback to the user.

[1852] Image data acquisition and analysis

[1853] When a user wants to check a specific place or object, they capture an image of it through the smart glasses. The device sends the captured image data to a server. The server analyzes the received image data and generates a description of the place or the name of the object. For example, if a user is looking for a conference room, the server generates information such as "The conference room is just ahead on your right." The server sends the generated information to the device, which then provides it to the user.

[1854] Real-time question answering

[1855] The user asks a specific question by voice or text. For example, "What does this word mean?", the device sends the question data to the server. The server analyzes the question data and generates the most appropriate answer. For example, it generates an answer such as "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[1856] Specific examples

[1857] Example 1: Use during a business meeting

[1858] If a user is unsure how to speak to a customer during a business meeting, the device will capture voice data in real time and display polite language suggestions such as "Thank you for taking the time."

[1859] Example 2: Using navigation

[1860] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is just ahead on your right."

[1861] Prompt Sentence Examples

[1862] Use the following prompt:

[1863] User Input: "Nice to meet you."

[1864] AI Model: "It's a pleasure to meet you."

[1865] User Input: "What does this word mean?"

[1866] AI Model: "This word means 'Tango'"

[1867] In this way, the system of the present invention allows users to receive appropriate communication support in real time, reducing difficulties in business and daily life. Furthermore, by using a generative AI model, it is possible to respond quickly and accurately to user needs, making the system useful for many people, regardless of age or level of technological literacy.

[1868] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1869] Acquisition and analysis of audio data

[1870] Step 1: The user puts on the smart glasses or smartwatch and starts a conversation.

[1871] Input: User's voice

[1872] Action: The user speaks.

[1873] Output: Audio signal

[1874] Step 2: The terminal acquires the user's voice data in real time.

[1875] Input: Audio signal

[1876] How it works: The device's microphone picks up audio signals and converts them into digital audio data.

[1877] Output: Digital audio data

[1878] Step 3: The device sends the acquired voice data to the server.

[1879] Input: Digital audio data

[1880] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[1881] Output: Compressed and encrypted audio data

[1882] Step 4: The server analyzes the received voice data using the generative AI model.

[1883] Input: Compressed and encrypted audio data

[1884] How it works: The server decodes the audio data and uses a generative AI model (e.g., GPT-4) to convert the audio to text and perform analysis.

[1885] Output: Analysis results (suggestions for appropriate honorifics and words)

[1886] Step 5: The server sends the generated proposal to the device.

[1887] Input: Analysis results (suggestions for appropriate honorifics and words)

[1888] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[1889] Output: Compressed and encrypted proposal data

[1890] Step 6: The device provides visual or auditory feedback to the user.

[1891] Input: Compressed and encrypted proposal data

[1892] How it works: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[1893] Output: User feedback (visual / auditory information)

[1894] Image data acquisition and analysis

[1895] Step 1: When a user wants to see a specific place or object, they capture an image of it through their smart glasses.

[1896] Input: Visual information

[1897] Action: The user operates the smartglasses camera to capture an image.

[1898] Output: Image data

[1899] Step 2: The device sends the captured image data to the server.

[1900] Input: Image data

[1901] Operation: The terminal's communication module compresses and encrypts the image data and sends it to the server.

[1902] Output: Compressed and encrypted image data

[1903] Step 3: The server analyzes the received image data.

[1904] Input: Compressed and encrypted image data

[1905] How it works: The server decodes the data and uses image recognition software to analyze the image (e.g., recognize location descriptions and object names).

[1906] Output: Analysis results (location descriptions and names of things)

[1907] Step 4: The server sends the generated information to the terminal.

[1908] Input: Analysis results (location description and object names)

[1909] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[1910] Output: Compressed and encrypted information data

[1911] Step 5: The terminal provides the user with information visually or audibly.

[1912] Input: Compressed and encrypted information data

[1913] How it works: The device decodes the data and displays it on the screen or speaks the information over the speaker.

[1914] Output: Provide information to the user (visual / auditory information)

[1915] Real-time question answering

[1916] Step 1: The user asks a specific question via voice or text.

[1917] Input: Voice or text question

[1918] How it works: If the user asks a question by voice, the device's microphone picks up the sound, and if the user asks a question by text, the device's text input function is used.

[1919] Output: Digital audio data or text data

[1920] Step 2: The device sends the query data to the server.

[1921] Input: Digital audio or text data

[1922] How it works: The device's communications module compresses and encrypts the data before sending it to the server.

[1923] Output: Compressed and encrypted query data

[1924] Step 3: The server analyzes the question data and generates the best answer.

[1925] Input: Compressed and encrypted query data

[1926] How it works: The server decodes the data, uses a generative AI model to analyze the intent of the question, and generates an answer.

[1927] Output: Analysis results (best answer)

[1928] Step 4: The server sends the generated response to the terminal.

[1929] Input: Analysis results (best answer)

[1930] How it works: The server compresses and encrypts the analysis results and sends them to the device.

[1931] Output: Compressed and encrypted answer data

[1932] Step 5: The device provides the user with a voice or text message.

[1933] Input: Compressed and encrypted answer data

[1934] What it does: The device decodes the data and displays it on the display or reads it aloud through the speaker.

[1935] Output: Provides answers to users (visual / auditory information)

[1936] As described above, the system achieves efficient communication support for users through specific operations and data input / output at each step.

[1937] (Application example 1)

[1938] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1939] Conventional communication support systems allow users to receive suggestions for appropriate honorific expressions and words, but they have difficulty answering customer questions in real time or providing in-store navigation information when dealing with customers in brick-and-mortar stores. They also lack the ability to provide personalized services based on customer facial recognition or past purchase history. The present invention provides a new communication support system that streamlines customer service in brick-and-mortar stores and improves customer satisfaction.

[1940] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1941] In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, means for analyzing question data acquired from the user via the smart wearable device and generating an optimal answer using a generative AI model, means for visually or audibly feeding back the generated answer, and means for displaying in-store navigation information based on the generated information, thereby enabling immediate responses to customer questions in physical stores, in-store navigation, and personalized services.

[1942] "Voice data" is information about sounds generated by the user's vocalizations.

[1943] A "server" is a central processing unit that processes and provides various data via a network.

[1944] "Analysis" is the act of interpreting and understanding the content of acquired data to generate appropriate information.

[1945] A "generative AI model" is a trained algorithm that uses artificial intelligence technology to analyze data and generate appropriate suggestions and answers.

[1946] Honorific language is the use of words to show respect to the other person, and is part of language usage.

[1947] "Word suggestion" is the act of selecting appropriate words and providing them to the user.

[1948] A "smart wearable device" is a portable electronic device that is worn by a user.

[1949] "Question data" is data of statements or text that a user uses to ask for specific information.

[1950] "Visual feedback" is a method of providing information to users in a form that they can see.

[1951] "Auditory feedback" is a method of providing information to a user in a form that can be heard by ear.

[1952] "Navigation information" is information that provides directions and location information to a specific location.

[1953] "Feedback" is the act of returning information or results from a system to a user.

[1954] The present invention is a communication support system for improving the efficiency of customer service in brick-and-mortar stores and enhancing customer satisfaction. This system is composed of three main elements: a user, a terminal (a smart wearable device), and a server. Specific embodiments for implementing the present invention are described below.

[1955] Acquisition and analysis of audio data

[1956] The user wears a device such as smart glasses or a smartwatch and responds to customers. The device captures the user's voice data in real time and sends that data to a server. The server then analyzes the voice data using a generative AI model and generates appropriate honorific and word suggestions. The generated suggestions are then fed back to the user visually or audibly via the device.

[1957] Image data acquisition and analysis

[1958] The user acquires image data of the interior of a store and the products through the device. The device then sends the captured image data to the server. The server analyzes the image data and generates location descriptions and names of objects. For example, it can also provide navigation information such as "The conference room is further ahead on the right."

[1959] Real-time question answering

[1960] When a user receives a question from a customer, the device sends the question data as voice or text to the server. The server analyzes the question data using a generative AI model and generates the optimal answer. This answer is also fed back to the user via the device.

[1961] Specific examples from physical stores

[1962] Example 1: Customer Service

[1963] When a customer asks, "Tell me about this new product," the user's device sends the question to the server, which uses a generative AI model to generate an answer such as, "This new product uses the latest AI technology and has a user-friendly interface," and the device visually displays the answer to the user.

[1964] Example 2: In-store navigation

[1965] When a customer asks, "Where is the conference room?", the user captures an image of the surroundings through the smart glasses and sends it to the server. The server analyzes the image data and generates navigation information such as "The conference room is on the second floor, on the right," which is then visually presented to the user on the device.

[1966] Hardware and software used

[1967] Hardware:

[1968] Smart Glasses

[1969] Smartwatch

[1970] microphone

[1971] software:

[1972] Python

[1973] SpeechRecognition Library

[1974] PIL (Python Imaging Library)

[1975] Requests library

[1976] OpenAI API

[1977] Prompt Sentence Examples

[1978] Here are some examples of prompts you can give to your generative AI model:

[1979] Customer Question: What are the features of this product?

[1980] Generate the appropriate answer.

[1981]

[1982] Please describe the image. The image shows a conference room.

[1983] The present invention allows users in physical stores to smoothly handle customer inquiries, provide instant responses to questions, and provide in-store navigation, which is expected to improve the quality of service and increase customer satisfaction.

[1984] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1985] Step 1:

[1986] The user puts on the smart glasses or smartwatch and begins interacting with the customer.

[1987] Input: User's voice or image data

[1988] Output: Captured audio or image data

[1989] Specific operation: The device's microphone captures audio data in real time, and the camera captures image data.

[1990] Step 2:

[1991] The terminal transmits the acquired audio and image data to the server.

[1992] Input: Audio or image data captured by the device

[1993] Output: Audio or image data transferred to the server

[1994] Specific operation: The terminal transmits the collected audio or image data to the server via the network.

[1995] Step 3:

[1996] The server analyzes the audio data and generates appropriate honorifics and word suggestions.

[1997] Input: Audio data received by the server

[1998] Output: Generated honorifics and word suggestions

[1999] Specific operation: The server processes the speech data using a generative AI model to generate appropriate honorifics and word suggestions.

[2000] Step 4:

[2001] The server analyzes the image data and generates location descriptions and names of objects.

[2002] Input: Image data received by the server

[2003] Output: Generated place descriptions and object names

[2004] How it works: The server processes image data using a generative AI model to identify location descriptions and object names.

[2005] Step 5:

[2006] The server sends generated suggestions and information to the terminal, which then provides visual or auditory feedback to the user.

[2007] Input: Generated suggestions and information from the server

[2008] Output: Feedback information provided on the device by display or audio

[2009] Specific operation: The server sends the generated honorific suggestions and navigation information to the device, which then displays them on the screen or notifies the user by voice.

[2010] Step 6:

[2011] It analyzes question data obtained from users via smart wearable devices and generates optimal answers using a generative AI model.

[2012] Input: Question data from the user

[2013] Output: The generated answer

[2014] How it works: The device receives a question via voice or text and sends it to the server, which then uses a generative AI model to analyze the question data and generate an answer.

[2015] Step 7:

[2016] The server generates a response and sends it to the terminal, which then provides visual or auditory feedback to the user.

[2017] Input: Generated answer from the server

[2018] Output: Answer information displayed on the device or provided as audio

[2019] Specific operation: The answer generated by the server is sent to the terminal and presented in a format that is easy for the user to see or hear.

[2020] Step 8:

[2021] The server displays in-store navigation information based on the generated information.

[2022] Input: Location information and image data from the user

[2023] Output: Generated navigation information

[2024] Specific operation: The server generates in-store navigation information based on location information and image data and sends it to the device, which then visually guides the user along the way.

[2025] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2026] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this embodiment is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[2027] System configuration

[2028] User

[2029] This is the subject who uses the system to receive communication support. They wear devices such as smart glasses or smart watches.

[2030] Terminal

[2031] A device that captures audio and image data in real time. In this case, it refers to smart glasses and smart watches. The device is responsible for sending the captured data to a server.

[2032] server

[2033] This is a central processing unit that analyzes the received voice and image data. It uses generative AI models to analyze the data and generate appropriate suggestions, which are then sent back to the device.

[2034] Emotion Engine

[2035] This engine recognizes the user's emotions from their voice and image data. Based on the emotion data, the server generates appropriate suggestions.

[2036] Program processing

[2037] Acquisition and analysis of voice and emotion data

[2038] 1. The user puts on the smart glasses or smartwatch and turns on the device.

[2039] 2. The device activates its audio and image sensors, and the user initiates a conversation.

[2040] 3. The terminal acquires the user's voice and image data in real time.

[2041] 4. The terminal sends the acquired voice and image data to the server.

[2042] 5. The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[2043] 6. The server generates appropriate suggestions based on the emotion engine analysis, such as "Great work. What would you like to talk about?"

[2044] 7. The server sends the generated suggestions to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (sound presented through earphones).

[2045] Image data acquisition and analysis

[2046] 1. When a user wants to check a specific place or object, they capture its image data through smart glasses.

[2047] 2. The device sends the captured image data to the server.

[2048] 3. The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expression. For example, if a user looks anxious while searching for a conference room.

[2049] 4. The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[2050] 5. The server sends the generated information to the terminal, which provides it to the user.

[2051] Real-time question answering

[2052] 1. The user speaks or writes a specific question, for example, "What does this word mean?"

[2053] 2. The device sends the question data to the server.

[2054] 3. The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expressions. For example, if the user is in a hurry.

[2055] 4. The server generates a quick and detailed answer based on the urgency of the question, such as "The meaning of this word is 'Tango'."

[2056] 5. The server sends the generated answer to the device, which provides it to the user via voice or text.

[2057] Specific examples

[2058] Example 1: Use during a business meeting

[2059] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[2060] Example 2: Using navigation

[2061] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[2062] As described above, the present invention allows users to receive appropriate communication support in real time, thereby reducing difficulties in business and daily life. This system is available to a wide range of users, regardless of age or level of understanding of technology, and is useful to many people.

[2063] The processing flow will be explained below.

[2064] Step 1:

[2065] The user puts on the smart glasses or smart watch and turns on the device. The device starts up and connects to the network.

[2066] Step 2:

[2067] The device activates its audio and image sensors and prepares to capture audio and images in real time.

[2068] Step 3:

[2069] When a user starts a conversation and has difficulty choosing words or using honorific language, they speak out verbally.

[2070] Step 4:

[2071] The terminal acquires the user's voice data in real time.

[2072] Step 5:

[2073] The terminal transmits the acquired voice data to the server.

[2074] Step 6:

[2075] The server analyzes the received voice data using a generative AI model.

[2076] Step 7:

[2077] The emotion engine recognizes the user's emotion based on the voice data. For example, if the user says, "I'm sorry, I'm in trouble," the emotion engine recognizes the user's confusion.

[2078] Step 8:

[2079] The server generates appropriate honorifics and word suggestions based on the analysis results of the emotion engine. For example, it generates a suggestion such as, "You seem to be in trouble. What can I explain to you?"

[2080] Step 9:

[2081] The server sends the generated proposal to the terminal.

[2082] Step 10:

[2083] The device provides the user with feedback on the analysis results either visually (displayed on the smart glasses display) or audibly (presented via earphones).

[2084] Step 11:

[2085] When a user wants to view a particular object or location, they capture image data of it through the smart glasses.

[2086] Step 12:

[2087] The terminal transmits the captured image data to the server.

[2088] Step 13:

[2089] The server analyzes the image data.

[2090] Step 14:

[2091] The emotion engine recognizes the user's emotional state from their facial expressions, for example, when the user has an anxious expression.

[2092] Step 15:

[2093] The server generates information including a description of the location and the name of the object based on the image analysis results and the recognition results of the emotion engine. For example, it could generate information such as "This is the conference room. It's up ahead on your right," providing additional information to ease the user's anxiety.

[2094] Step 16:

[2095] The server transmits the generated information to the terminal.

[2096] Step 17:

[2097] The device provides navigation information to the user visually (on the smart glasses display).

[2098] Step 18:

[2099] The user speaks or texts a specific question.

[2100] Step 19:

[2101] The terminal transmits the question data to the server.

[2102] Step 20:

[2103] The server parses the question data.

[2104] Step 21:

[2105] The emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression when asking the question. For example, if it determines that the question is urgent,

[2106] Step 22:

[2107] The server generates a quick and detailed answer based on the evaluation of the emotion engine, for example, "The meaning of this word is 'Tango'."

[2108] Step 23:

[2109] The server generates a response and sends it to the terminal.

[2110] Step 24:

[2111] The device provides the user with a response by voice or text.

[2112] Step 25:

[2113] The user selects a subscription plan using a dedicated application.

[2114] Step 26:

[2115] The terminal transmits the selected plan information to the server.

[2116] Step 27:

[2117] The server enables certain features based on the plan selected by the user.

[2118] Step 28:

[2119] The server monitors the user's usage and sends suggestions and notifications for plan changes to the device as needed.

[2120] Through this process, the system enables users to receive appropriate communication assistance in real time, easing difficulties in business and daily life.

[2121] Example 2

[2122] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2123] Conventional communication support systems have had difficulty in providing real-time support by fully utilizing the user's voice data and image data. In particular, they lacked the ability to provide appropriate suggestions and information that took the user's emotions into consideration, making it difficult to improve the user experience. They also lacked the ability to provide prompt and appropriate answers to user questions. It is necessary to provide a system that can solve these problems and enable users to receive better support.

[2124] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2125] In this invention, the server includes means for acquiring and processing a user's voice data and image data in real time, means for transmitting the voice data and image data to the server, means for analyzing the voice data and image data by the server and recognizing emotions, means for analyzing the voice data using a generative AI model to generate appropriate suggestions, and means for feeding back the generated suggestions to the user. This allows the user to receive appropriate suggestions and information that take emotions into consideration in real time, significantly improving the communication support experience. It also enables the provision of quick and appropriate answers to questions.

[2126] A "user" is an entity that utilizes the communication support system to provide voice and image data and receive support.

[2127] A "terminal" is a device that captures audio and image data in real time and transmits it to a server, and examples of this include smart glasses and smart watches.

[2128] The "server" is a central processing unit that analyzes the received audio and image data, generates appropriate suggestions, and sends them to the terminal.

[2129] The "emotion engine" is an analysis engine for recognizing the user's emotions from the user's voice data and image data.

[2130] A "generative AI model" is an artificial intelligence model that runs on a server and analyzes voice and image data to generate appropriate suggestions and answers.

[2131] "Voice data" refers to data that is a digital recording of a user's speech.

[2132] "Image data" is digital data that captures the user's face and environment with a camera.

[2133] "Analysis" is the process of extracting information from the acquired voice data and image data and understanding their meaning.

[2134] "Suggestions" are advice or information for the user that are generated by the server based on the results of the analysis.

[2135] "Question data" is voice or text data containing a specific question from the user.

[2136] An "answer" is a solution generated by the server by analyzing question data, and is information provided to the user.

[2137] "Feedback" is the act of providing suggestions or answers to a user, and can be done visually or audibly.

[2138] This invention combines an emotion engine with an AI-based communication support system to provide communication support by analyzing the user's voice and image data in real time. The system of this invention is composed of four main elements: the user, the terminal, the server, and the emotion engine.

[2139] First, the user puts on a device such as smart glasses or a smart watch and turns it on. The device activates its audio and image sensors, and the user begins a conversation. The device captures the user's audio and image data in real time and sends the acquired data to a server. The device can be a common wearable device such as smart glasses or a smart watch. These devices communicate with the server via Wi-Fi or Bluetooth.

[2140] Next, the server analyzes the received voice data using a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to talk about," the emotion engine recognizes that the user is confused. Based on the emotion engine's analysis results, the server generates an appropriate suggestion, such as, "Thank you for your hard work. What would you like to talk about?" The server then sends the generated suggestion to the device, which then provides feedback to the user visually (displayed on the smart glasses display) or audibly (audio provided through earphones).

[2141] Furthermore, if a user wants to check a specific location or object, they can capture image data of that location through the smart glasses. The device sends the captured image data to a server, which then analyzes the image data. This analysis includes cases where the user has an anxious expression, such as when searching for a conference room. The server generates navigation information such as "The conference room is further ahead on your right," and provides additional information that takes the user's emotions into consideration. The server then sends the generated information to the device, which then provides it to the user.

[2142] In addition, a real-time question-and-answer function is also provided. The user inputs a specific question by voice or text, and the device sends the question data to the server. The server analyzes the question data, and an emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression. For example, if the user is in a hurry, the server will take the urgency into account and generate a quick and detailed answer such as, "The meaning of this word is 'Tango'." The server then sends the generated answer to the device, which then provides it to the user by voice or text.

[2143] Specific examples

[2144] Example 1: Use during a business meeting

[2145] When a user is unsure how to speak to a customer during a business meeting, the device collects voice data in real time and suggests polite expressions such as, "Thank you for your time." In this case, the emotion engine recognizes the user's nervousness and selects expressions that will calm them down.

[2146] Example 2: Using navigation

[2147] When a user is searching for a conference room in a large office building, the smart glasses capture images of the surrounding area. The server analyzes the image data and provides navigation information such as, "The conference room is further ahead on your right." At the same time, the emotion engine recognizes the user's anxiety and adds encouraging expressions.

[2148] Prompt Sentence Examples

[2149] Below are some examples of prompts to input to the generative AI model:

[2150] "Please analyze the audio data that the user is confused about."

[2151] "Analyze image data showing a user's anxious expression."

[2152] "Generate quick answers to user questions"

[2153] As a result, users can receive appropriate communication support in real time, and difficulties in business and daily life can be alleviated.

[2154] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2155] Step 1:

[2156] The user puts on the smart glasses or smartwatch and turns on the device.

[2157] Input: A user presses the power button on the device.

[2158] Operation: The device boots up and performs self-tests and initialization.

[2159] Output: The device displays a status of ready.

[2160] Step 2:

[2161] The device activates its audio and image sensors, and the user initiates a conversation.

[2162] Input: Device readiness status and user voice input.

[2163] Action: The device activates the audio sensors and camera and begins capturing data.

[2164] Output: Audio and image data captured by the device.

[2165] Step 3:

[2166] The terminal acquires the user's voice and image data in real time.

[2167] Input: User's speech and facial expressions.

[2168] How it works: The audio sensor converts sound into digital data, and the camera captures images.

[2169] Output: Acquired audio and image data.

[2170] Step 4:

[2171] The terminal transmits the acquired voice data and image data to the server.

[2172] Input: Captured audio and image data.

[2173] How it works: Your device sends data to a server via Wi-Fi or Bluetooth.

[2174] Output: Audio and image data sent to the server.

[2175] Step 5:

[2176] The server analyzes the voice data using a generative AI model, and the emotion engine analyzes the user's emotions.

[2177] Input: Transmitted audio and image data.

[2178] How it works: The server inputs voice data into a generative AI model, and the emotion engine analyzes the user's emotions. For example, if a user says, "I'm sorry, I don't know what to say," the emotion engine recognizes that the user is confused.

[2179] Output: Parsed sentiment data and appropriate suggestions.

[2180] Step 6:

[2181] The server generates appropriate suggestions based on the analysis results of the emotion engine, such as "Good work. What would you like to talk about?"

[2182] Input: Parsed emotion data and the analysis results of the generative AI model.

[2183] How it works: The server generates appropriate suggestions and converts them into text or audio format.

[2184] Output: The generated proposals.

[2185] Step 7:

[2186] The server sends the generated suggestions to the terminal, which provides visual or auditory feedback to the user.

[2187] Input: Generated proposals.

[2188] How it works: The server sends the suggestions to the device, which then provides feedback to the user, for example, on the smart glasses display or as audio through earphones.

[2189] Output: Feedback provided to the user.

[2190] Step 8:

[2191] When a user wants to view a particular place or object, they capture image data of it through the smart glasses.

[2192] Input: The situation in which you are looking at the place or thing you want to check.

[2193] How it works: The user operates the smart glasses and the camera captures image data.

[2194] Output: The captured image data.

[2195] Step 9:

[2196] The terminal transmits the captured image data to the server.

[2197] Input: The captured image data.

[2198] Operation: The device compresses the image data and sends it to the server.

[2199] Output: Image data sent to the server.

[2200] Step 10:

[2201] The server analyzes the image data, and the emotion engine recognizes anxiety or confusion from the user's facial expressions.

[2202] Input: The submitted image data.

[2203] How it works: The server analyzes the image data, and the emotion engine recognizes facial expressions. For example, it analyzes the facial expressions of a user searching for a meeting room.

[2204] Output: Parsed emotion data and pertinent information.

[2205] Step 11:

[2206] The server generates navigation information such as "The conference room is further ahead on the right," and provides additional information that takes the user's feelings into consideration.

[2207] Input: Parsed emotion data and user location information.

[2208] How it works: The server generates navigation information and emotion-sensitive messages.

[2209] Output: Generated navigation information and additional information.

[2210] Step 12:

[2211] The server transmits the generated information to the terminal, which then provides it to the user.

[2212] Input: Generated navigation information and additional information.

[2213] How it works: The server sends information to the device, which then provides it to the user, for example by displaying it on smart glasses or providing audio guidance.

[2214] Output: Information provided to the user.

[2215] Step 13:

[2216] The user speaks or texts a specific question.

[2217] Input: The user's question.

[2218] How it works: A user types a question using the microphone or text input function on their smart glasses.

[2219] Output: The input question data.

[2220] Step 14:

[2221] The terminal transmits the question data to the server.

[2222] Input: The entered question data.

[2223] Operation: The device sends the query data to the server.

[2224] Output: The query data sent to the server.

[2225] Step 15:

[2226] The server analyzes the question data, and the emotion engine evaluates the urgency and importance of the question based on the user's voice and facial expression.

[2227] Input: Submitted question data and user sentiment data.

[2228] How it works: The server analyzes the question data, and the emotion engine evaluates the urgency and importance.

[2229] Output: Assessed urgency and importance.

[2230] Step 16:

[2231] The server takes into account the urgency and generates a quick and detailed response.

[2232] Input: Assessed urgency and question data.

[2233] How it works: The server uses a generative AI model to generate an answer, such as "The meaning of this word is 'Tango'."

[2234] Output: The generated answer.

[2235] Step 17:

[2236] The server generates a response and sends it to the terminal, which provides it to the user in voice or text.

[2237] Input: The generated answer.

[2238] Operation: The server sends the answer data to the terminal, which then provides it to the user via voice or text.

[2239] Output: The answer provided to the user.

[2240] (Application example 2)

[2241] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2242] Current communication support systems can make appropriate suggestions using the user's voice and image data, but they cannot generate responses based on the user's emotions, making it difficult to provide prompt and appropriate support.In addition, in certain environments such as factories, there are no systems that can properly analyze and respond to workers' difficulties and stress, making it difficult to improve productivity.

[2243] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice data in real time, means for transmitting the voice data to the server, means for analyzing the voice data by the server and generating appropriate honorific expressions and word suggestions, means for feeding back the generated suggestions to the user, and means including an emotion engine for analyzing the user's emotion data and generating responses. This makes it possible to provide the user with appropriate suggestions and responses based on the emotion data, thereby achieving more effective communication support.

[2244] A "user" is an entity that uses the system to receive communication support.

[2245] "Voice data" refers to data that records the voice uttered by the user.

[2246] A "server" is a central processing unit that analyzes the data received and generates the necessary suggestions and responses.

[2247] "Emotion data" refers to data that indicates emotional information extracted from the user's voice or image.

[2248] An "emotion engine" is a system component that analyzes emotion data and recognizes the user's emotions.

[2249] A "generative AI model" is an artificial intelligence model that generates appropriate responses and suggestions based on input data.

[2250] "Suggestion" refers to the generation of useful information or advice for the user.

[2251] "Feedback" is the act of providing server-generated suggestions and responses to the user.

[2252] "Question data" is data including the content of a question provided by a user.

[2253] This invention provides a system for an emotionally responsive communication robot that supports workers in factories. Specifically, the robot acquires the worker's voice data and image data in real time, sends this data to a server for analysis, and generates appropriate suggestions and responses based on the worker's emotions. This reduces worker stress and helps improve productivity.

[2254] Hardware and software used

[2255] Hardware: A robotic device equipped with audio and image sensors, a server as the central processing unit.

[2256] Software: Speech analysis software (Google Cloud Speech-to-Text API), image analysis software (OpenCV), emotion engine (Microsoft Azure Emotion API), generative AI model (OpenAI GPT-4).

[2257] Program processing

[2258] 1. Acquisition and analysis of voice and emotion data

[2259] The worker begins working in front of the robot.

[2260] The robot activates audio and image sensors to capture the worker's voice and facial expressions in real time.

[2261] The robot transmits the acquired voice data and image data to the server.

[2262] The server analyzes the voice data using the Google Cloud Speech-to-Text API and the emotion data using the Microsoft Azure Emotion API.

[2263] Based on the emotional data, suggestions tailored to the worker's situation are generated using a generative AI model (OpenAI GPT-4).

[2264] The server sends the generated suggestions back to the robot, which then provides the suggestions to the worker via voice or display.

[2265] 2. Image Data Acquisition and Analysis

[2266] If a worker is having difficulty with a particular task, the robot will capture that situation.

[2267] The robot transmits the captured image data to a server.

[2268] The server uses OpenCV to analyze the image data, and the emotion engine recognizes the emotions from the worker's facial expressions.

[2269] Appropriate suggestions based on the worker's emotions are generated by a generative AI model, and then communicated by the robot.

[2270] Specific examples

[2271] 1. Support for workers

[2272] When a worker is performing a difficult task, the robot detects the "difficulty" from the worker's voice and facial expression.

[2273] The robot will make suggestions such as, "Thank you for your hard work. Is there anything I can help you with?"

[2274] This allows workers to receive appropriate support and carry out their work efficiently.

[2275] 2. Suggest short breaks

[2276] If a worker looks tired, the emotion engine will recognize this.

[2277] An example of a prompt from a generative AI model: "You seem to be tired. How about taking a short break?" Based on this, the model suggests, "How about taking a short break?"

[2278] This reduces worker fatigue and improves productivity.

[2279] In this way, communication support can be realized by providing appropriate suggestions and responses based on the user's emotions. This system is also very useful for supporting workers in factories and other specific environments.

[2280] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2281] Step 1:

[2282] When the user starts working in front of the robot, the robot activates the audio and image sensors. The audio sensor then captures the user's voice data, and the image sensor captures the user's facial expressions. Both data (audio and image data) are acquired in real time.

[2283] Input: User's voice and image

[2284] Output: Captured audio and image data

[2285] Step 2:

[2286] The device sends the acquired audio and image data to the server, where it is ready to be analyzed.

[2287] Input: Acquired audio and image data

[2288] Output: Audio and image data sent to the server

[2289] Step 3:

[2290] The server converts the voice data into text using the Google Cloud Speech-to-Text API, which is then used for analysis by the emotion engine. It also uses OpenCV to extract facial expression data from the image data.

[2291] Input: Audio and image data sent to the server

[2292] Output: Analyzed voice data (text data) and facial expression data

[2293] Step 4:

[2294] The emotion engine (Microsoft Azure Emotion API) analyzes the user's emotions from the voice and facial expression data. The emotion data indicates the user's current emotional state (e.g., confusion, fatigue, tension, etc.).

[2295] Input: Analyzed voice data and facial expression data

[2296] Output: User emotion data

[2297] Step 5:

[2298] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate suggestions and responses based on the emotion data. Specific prompts are provided to the generative AI model using the emotion data as input.

[2299] Input: User emotion data

[2300] Output: The generated suggestions and responses

[2301] Step 6:

[2302] The server sends the generated suggestions and responses to the robot, which then provides the suggestions and responses to the user through voice or a display. For example, a message such as "Good work! Is there anything I can help you with?" may be displayed or spoken.

[2303] Input: Generated suggestions and responses

[2304] Output: The suggestions or responses provided to the user

[2305] Step 7:

[2306] The user receives suggestions and responses from the robot and, if necessary, further communicates with the robot. The user's reactions and additional audio and video data are processed again starting from step 1.

[2307] Input: User reactions to the suggestions and responses provided

[2308] Output: New audio and image data for the next cycle

[2309] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2310] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2311] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2312] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2313] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2314] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2315] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2316] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2317] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2318] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2319] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2320] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2321] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2322] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2323] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2324] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2325] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2326] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2327] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2328] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2329] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2330] The following is further disclosed regarding the above embodiment.

[2331] (Claim 1)

[2332] means for acquiring user voice data in real time;

[2333] means for transmitting the voice data to a server;

[2334] A means for analyzing the speech data by the server and generating appropriate honorific expressions and word suggestions;

[2335] means for feeding back the generated suggestions to a user;

[2336] A communication support system including:

[2337] (Claim 2)

[2338] means for acquiring user image data;

[2339] means for transmitting the image data to a server;

[2340] means for analyzing image data by the server and generating a description of a place or a name of an object;

[2341] 2. The communication support system according to claim 1, further comprising: means for providing the generated information to a user.

[2342] (Claim 3)

[2343] A means for acquiring question data from a user;

[2344] means for transmitting the question data to a server;

[2345] means for analyzing question data by the server and generating an answer;

[2346] 2. The communication support system according to claim 1, further comprising means for providing the generated answer to the user.

[2347] "Example 1"

[2348] (Claim 1)

[2349] means for acquiring user voice data in real time;

[2350] means for transmitting said audio data to a central processing unit;

[2351] means for analyzing the speech data by the central processing unit and generating appropriate honorific expressions and word suggestions;

[2352] means for feeding back the generated suggestions to a user;

[2353] means including a generative AI model for analyzing the audio data;

[2354] A system including:

[2355] (Claim 2)

[2356] means for acquiring user image data;

[2357] means for transmitting the image data to a central processing unit;

[2358] means for analyzing image data by said central processing unit to generate location descriptions and object names;

[2359] means for providing the generated information to a user;

[2360] 10. The system of claim 1, further comprising means including image recognition software for analyzing said image data.

[2361] (Claim 3)

[2362] A means for acquiring question data from a user;

[2363] means for transmitting the query data to a central processing unit;

[2364] means for analyzing question data and generating answers by the central processing unit;

[2365] means for providing the generated answer to a user;

[2366] 10. The system of claim 1, further comprising: means for including a generative AI model for analyzing the query data.

[2367] "Application Example 1"

[2368] (Claim 1)

[2369] means for acquiring user voice data in real time;

[2370] means for transmitting the voice data to a server;

[2371] A means for analyzing the speech data by the server and generating appropriate honorific expressions and word suggestions;

[2372] means for feeding back the generated suggestions to a user;

[2373] A means for analyzing question data acquired from a user via a smart wearable device and generating an optimal answer using a generative AI model;

[2374] means for providing visual or auditory feedback of the generated answer;

[2375] a means for displaying in-store navigation information based on the generated information;

[2376] A system including:

[2377] (Claim 2)

[2378] means for acquiring user image data;

[2379] means for transmitting the image data to a server;

[2380] means for analyzing image data by the server and generating a description of a place or a name of an object;

[2381] The system of claim 1 further comprising: means for providing the generated information to a user.

[2382] (Claim 3)

[2383] A means for acquiring question data from a user;

[2384] means for transmitting the question data to a server;

[2385] means for analyzing question data by the server and generating an answer;

[2386] The system of claim 1 further comprising means for providing the generated answer to a user.

[2387] "Example 2: Combining Emotion Engines"

[2388] (Claim 1)

[2389] means for acquiring and processing user voice data and image data in real time;

[2390] means for transmitting the audio data and image data to a server;

[2391] means for analyzing voice data and image data by the server and recognizing emotions;

[2392] a means for analyzing the speech data using a generative AI model to generate appropriate suggestions;

[2393] means for feeding back the generated suggestions to a user;

[2394] means for feeding back the generated suggestions to a user;

[2395] A communication support system including:

[2396] (Claim 2)

[2397] means for acquiring user image data and transmitting it to a server;

[2398] means for analyzing image data by the server and generating information that takes into consideration a description of a place, a name of an object, and emotion based on emotion data of a user;

[2399] The system of claim 1 further comprising: means for providing the generated information to a user.

[2400] (Claim 3)

[2401] means for acquiring question data from a user and transmitting the data to a server;

[2402] means for analyzing question data by the server, evaluating the user's feelings, and generating an answer according to the degree of urgency;

[2403] The system of claim 1 further comprising means for providing the generated answer to a user.

[2404] "Application example 2 when combining emotion engines"

[2405] (Claim 1)

[2406] means for acquiring user voice data in real time;

[2407] means for transmitting the voice data to a server;

[2408] A means for analyzing the speech data by the server and generating appropriate honorific expressions and word suggestions;

[2409] means for feeding back the generated suggestions to a user;

[2410] means including an emotion engine for analyzing the user's emotion data and generating a response;

[2411] A communication support system including:

[2412] (Claim 2)

[2413] means for acquiring user image data;

[2414] means for transmitting the image data to a server;

[2415] means for analyzing image data by the server and generating a description of a place or a name of an object;

[2416] means for providing the generated information to a user;

[2417] 2. The communication support system according to claim 1, further comprising an emotion engine that analyzes emotion data of a user and generates a response.

[2418] (Claim 3)

[2419] A means for acquiring question data from a user;

[2420] means for transmitting the question data to a server;

[2421] means for analyzing question data by the server and generating an answer;

[2422] means for providing the generated answer to a user;

[2423] 2. The communication support system according to claim 1, further comprising means for optimizing the answers generated by the generative AI model in conjunction with the user's emotional data. [Explanation of symbols]

[2424] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for acquiring user voice data in real time; means for transmitting the voice data to a server; A means for analyzing the speech data by the server and generating appropriate honorific expressions and word suggestions; means for feeding back the generated suggestions to a user; A communication support system including:

2. means for acquiring user image data; means for transmitting the image data to a server; means for analyzing image data by the server and generating a description of a place or a name of an object; 2. The communication support system according to claim 1, further comprising means for providing the generated information to a user.

3. A means for acquiring question data from a user; means for transmitting the question data to a server; means for analyzing question data by the server and generating an answer; 2. The communication support system according to claim 1, further comprising means for providing the generated answer to the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A