System

An earphone-type device with real-time conversational support using natural language processing and AI generates appropriate phrases, addressing the challenge of instant phrase generation in critical communication scenarios.

JP2026019737APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121485
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Users face challenges in generating appropriate and humorous conversational phrases in real-time, particularly in industries like public relations, sales, and education, where smooth communication is crucial.

Method used

An earphone-type device collects conversations in real-time, transmits audio data to a server for analysis using natural language processing and artificial intelligence, generates appropriate phrases, and delivers them back to the user as whispers through the earphone speaker.

Benefits of technology

Enables users to engage in smooth and engaging conversations by providing timely and humorous phrases, enhancing communication skills in various situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019737000001_ABST
    Figure 2026019737000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting a conversation of a user; means for transmitting the collected conversation to a server; means for analyzing the conversation and generating an appropriate phrase in the server; means for transmitting the generated phrase to a terminal; and means for reproducing the phrase in a whisper in the terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional communication, users have difficulty instantly coming up with appropriate remarks or humorous phrases. This problem is particularly pronounced in industries where communication is important, such as public relations, sales, and education, where users experience difficulties in situations where smooth and effective conversation is required. Technology that can instantly provide appropriate conversational phrases is needed. [Means for solving the problem]

[0005] The present invention provides an earphone-type device equipped with a means for collecting a user's conversation in real time and transmitting it to a server. The server analyzes the received conversation using natural language processing technology and generates appropriate phrases. Furthermore, the system provides a system that transmits the generated phrases to a terminal, which then suggests them to the user in a whisper using speech synthesis technology, allowing the user to lead a smooth and engaging conversation. By using mood analysis means and artificial intelligence algorithms, the accuracy of phrase generation is improved, and real-time, low-latency communication enables the system to provide immediate feedback.

[0006] The "means for collecting the user's conversation" is a function for capturing the conversation of the user and those around them in real time using a microphone built into the earphone-type device and related technology.

[0007] The "means for transmitting collected conversations to a server" is a function that appropriately compresses and encrypts the voice data collected by the terminal and transmits it to the server with low latency.

[0008] "Means for analyzing the conversation on the server and generating appropriate phrases" refers to the process in which the server receives the voice data, analyzes the context and mood of the conversation using natural language processing technology and artificial intelligence, and generates appropriate phrases based on that analysis.

[0009] The "means for transmitting the generated phrase to the terminal" is a function for the server to quickly transmit the generated phrase to the terminal, thereby providing suggestions to the user in real time.

[0010] "Means for playing phrases in a whisper on the terminal" refers to a function that converts received phrases into voice using speech synthesis technology and delivers them to the user as a whisper through the earphone speaker.

[0011] "Natural language processing technology" is a computer science technology for analyzing voice and text data and understanding, classifying, and generating its content.

[0012] An "artificial intelligence algorithm" is a computational method that uses technologies such as machine learning and deep learning to automatically learn patterns from data and make inferences and predictions.

[0013] "Mood analysis means" is a technique that determines the tone and emotional state of a conversation and analyzes the mood of the conversation based on that.

[0014] A "low latency communication protocol" is a communication method that minimizes the delay time it takes for data to travel from the sender to the server where the necessary analysis is performed, and then back to the receiver.

[0015] "Speech synthesis technology" is a technology for converting text data into voice data that resembles the human voice. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, transmits the audio data to a server for analysis, and sends appropriate phrases generated by the analysis back to the terminal, where they are "whispered" to suggest the phrases to the user. The present invention is implemented as follows.

[0038] 1. Recording conversations using a device

[0039] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory. Noise cancellation processing is applied to obtain clear audio. The device prepares to send this audio data to a server at regular intervals.

[0040] 2. Sending conversations by device

[0041] The device compresses the voice data and sends it to the server as data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted to ensure security. Data transmission is always in real time, with great care taken to minimize latency.

[0042] 3. Analysis of conversations by the server

[0043] The server decodes the received audio data and converts it into text using natural language processing (NLP). The converted text undergoes contextual analysis and keyword extraction to identify important keywords and topics. Mood analysis is also performed to evaluate the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[0044] 4. Phrase generation by the server

[0045] The server generates appropriate and humorous phrases based on the extracted keywords and mood information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. This algorithm is capable of creating natural phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and immediately sent to the device.

[0046] 5. Proposal transmission from the server to the device

[0047] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[0048] 6. "Whisper" phrases from your device

[0049] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[0050] Specific examples

[0051] Example 1: Business meeting

[0052] When a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts the keywords "project" and "want to talk," performs a mood analysis, and generates an appropriate phrase, such as "Why don't you talk a little about the progress of the project?" The generated phrase is sent to the device and played as a "whisper" in the user's ear.

[0053] Example 2: Casual conversation

[0054] When a user says, "Do you have any plans for the weekend?", the device collects the corresponding voice and sends it to the server. The server analyzes "weekend" and "plans" as keywords and also considers the user's emotional state (e.g., interests), generating the phrase "Why don't you invite me to a new restaurant that's been making waves lately?". This phrase is then whispered through the device: "Do you want to try a restaurant that's been making waves lately?"

[0055] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations.

[0056] The processing flow will be explained below.

[0057] Step 1:

[0058] The user wears the earphone-type device. The device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory.

[0059] Step 2:

[0060] The device applies noise cancellation to the collected voice data to make it clear. The voice data is sampled at regular intervals, compressed into data packets, and encrypted.

[0061] Step 3:

[0062] The device sends compressed and encrypted audio data to the server using a low-latency communication protocol to maintain real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[0063] Step 4:

[0064] The server decodes the received voice data and converts it into text using natural language processing (NLP) techniques.

[0065] Step 5:

[0066] The server performs contextual analysis and keyword extraction on the converted conversation content, identifying important keywords and topics and understanding the context of the conversation.

[0067] Step 6:

[0068] The server performs mood analysis to assess the tone and emotional state of the conversation, thereby capturing the atmosphere of the conversation (e.g., joy, tension, excitement).

[0069] Step 7:

[0070] The server generates appropriate and humorous phrases based on context and mood information, using pre-trained artificial intelligence (AI) algorithms.

[0071] Step 8:

[0072] The server then re-compresses the generated phrases into data packets, which are then sent to the device using a low-latency communication protocol.

[0073] Step 9:

[0074] The device decompresses the received phrase data and converts it from text to voice using text-to-speech technology.

[0075] Step 10:

[0076] The device then plays the generated voice data through the earphone speaker in a whisper close to the user's ear, allowing the user to continue the conversation using phrases suggested in real time.

[0077] Step 11:

[0078] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data back to the server.

[0079] Step 12:

[0080] The server feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm.

[0081] Example 1

[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0083] In many modern conversational environments, users face the challenge of generating appropriate phrases in real time and smoothly advancing a conversation. In particular, in business meetings and casual conversations, being able to instantly respond appropriately to the situation requires high communication skills. Therefore, there is a need for technology that allows users to obtain appropriate phrases in a timely manner during a conversation.

[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0085] In this invention, the server includes a means for analyzing speech and converting it into text data, a means for using a generative AI model to generate appropriate phrases based on the text data, and a means for analyzing the user's emotional state, thereby enabling the server to analyze the user's conversation in real time and instantly generate and provide appropriate and humorous phrases.

[0086] "User" refers to an individual who wears an earphone-type device and is the conversation partner.

[0087] "Voice" refers to conversations and other vocal content spoken by a user.

[0088] "Collection means" refers to the equipment and technology used to pick up and appropriately process audio.

[0089] A "terminal" is an earphone-type device worn by a user, and refers to a device that collects, processes, and transmits audio.

[0090] "Server" refers to a computer system for receiving voice data, analyzing it, converting it into text, generating phrases, etc.

[0091] "Generative AI model" refers to an algorithm equipped with artificial intelligence technology that is used to generate appropriate phrases based on text data.

[0092] "Audio playback means" refers to a technique or device that plays back the generated phrase as audio.

[0093] "Natural language processing technology" refers to computer technology for processing, understanding, and generating human language.

[0094] "Emotional state" refers to the user's psychological state, which indicates the tone and emotional state of the conversation.

[0095] The present invention relates to a system in which a user wears an earphone-type device, collects the sound of a conversation in real time, and transmits the collected sound data to a server for analysis. The system of the present invention is implemented as follows.

[0096] 1. The user puts on the device

[0097] When a user puts the earphone-type device in their ear, the device automatically starts up and turns on sound collection mode, allowing it to collect conversations around the user in real time.

[0098] 2. Recording conversations using a device

[0099] The device is equipped with a highly sensitive microphone that collects conversations around the user in real time. The audio data is temporarily stored in the device's local memory, and noise-cancelling technology is used to ensure clear audio. This process uses a microphone with noise-cancelling technology.

[0100] 3. Transmission of audio data by the device

[0101] The device compresses the collected audio data and sends it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end to ensure security. An audio compression codec such as G.711 is used for this compression.

[0102] 4. Receiving and analyzing audio data by the server

[0103] The server receives the voice data sent from the device and uses a speech recognition engine (e.g., Google Speech-to-Text) to decode it. The server converts the voice data into text data, then uses natural language processing (NLP) techniques to analyze the context of the voice and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state.

[0104] 5. Server-generated phrases

[0105] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. This generative AI model has the ability to predict and generate the most appropriate phrases based on the user's speech content and emotional state.

[0106] 6. Sending proposals from the server to the device

[0107] The generated phrase is then compressed again and sent to the device using a low-latency communication protocol. The device then decodes the compressed data and moves on to the next step.

[0108] 7. "Whisper" phrases from your device

[0109] The device converts the received phrase into a voice in real time using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear. This voice playback uses a TTS engine (e.g., Amazon Polly).

[0110] Specific examples

[0111] Example 1: Business meeting

[0112] User: "Today I'd like to talk about a new project."

[0113] The device collects the audio and sends it to the server.

[0114] The server analyzes keywords such as "project" and "want to talk" and performs mood analysis.

[0115] The server uses a "generative AI model" to generate the phrase "Why not talk a little about the progress of the project?" and sends it to the device.

[0116] The terminal whispers, "Why not give us a little update on the progress of the project?"

[0117] Example 2: Casual conversation

[0118] User: "Do you have any plans for the weekend?"

[0119] The device collects the audio and sends it to the server.

[0120] The server analyzes "weekend" and "schedule" and takes into account the user's emotional state.

[0121] The server uses a "generative AI model" to generate the phrase "Why not invite me to a new restaurant that's the talk of the town?" and sends it to the device.

[0122] The device whispers, "Would you like to try a restaurant that's been getting a lot of attention lately?"

[0123] In this way, the present invention provides users with appropriate and humorous phrases in real time, allowing for smooth and engaging conversations.

[0124] Example prompt sentence:

[0125] "When a user says, 'Today I'd like to talk about a new project' in a business meeting, please suggest an appropriate phrase."

[0126] Example output:

[0127] "Maybe we should talk a little bit about how the project is going."

[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0129] Step 1:

[0130] User Wears Device

[0131] The user puts the earphone-type device in their ear. This causes the device to automatically start up and turn on sound collection mode. The input is the state in which the earphone is being worn, and the output is the device switching to sound collection mode.

[0132] Step 2:

[0133] Recording conversations using a device

[0134] The device uses a built-in microphone to collect conversations around the user in real time. The collected audio data is temporarily stored in local memory, and clear audio is obtained using noise-canceling technology. The input is the surrounding audio, and the output is clear audio data that has been subjected to noise-canceling processing. Specifically, the device collects audio every second and stores the data in a buffer.

[0135] Step 3:

[0136] Sending audio data by the device

[0137] The device compresses the collected voice data using a voice compression codec such as G.711 and transmits it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end before transmission. The input is clear voice data, and the output is compressed and encrypted data packets.

[0138] Step 4:

[0139] Receiving and analyzing voice data by the server

[0140] The server receives and decodes the voice data sent from the device. It then uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. It then uses natural language processing (NLP) technology to analyze the context of the text and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state. The input is the decoded voice data, and the output is the analyzed text and emotional information. Specifically, the server receives the voice data, converts it into text through the speech recognition engine, and then analyzes it using an NLP model.

[0141] Step 5:

[0142] Server-generated phrases

[0143] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. The generative AI model predicts and generates the most appropriate phrase based on the prompt sentence. The input is the text data and emotional information, and the output is the generated phrase. Specifically, the server inputs the prompt sentence into the generative AI model and obtains the generated phrase.

[0144] Step 6:

[0145] Sending proposals from the server to the device

[0146] The server compresses the generated phrase and transmits it to the terminal using a low-latency communication protocol. The input is the generated phrase, and the output is a compressed and encrypted data packet. Specifically, the server compresses the generated phrase and encrypts it for transmission as a data packet.

[0147] Step 7:

[0148] "Whisper" phrases from your device

[0149] The device converts the received phrase into audio in real time using Text-to-Speech (TTS) technology and plays it as a "whisper" in the user's ear. The input is encrypted phrase data, and the output is a spoken suggested phrase. Specifically, the device uses a TTS engine to convert the phrase into audio and plays the audio through earphones.

[0150] (Application example 1)

[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0152] In conventional content distribution services, it has been difficult to provide appropriate commentary or supplemental information about the content being viewed in real time. It has also been a challenge to provide accurate information according to the user's conversation and mood. The present invention aims to solve these problems.

[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0154] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for reproducing the phrases in a whisper in the terminal, and means for analyzing the content being viewed and providing commentary and supplemental information in real time, thereby enabling the user to enjoy appropriate commentary and supplemental information in real time even while viewing the content.

[0155] The "means for collecting the user's conversation" is a function for collecting the user's voice using a microphone or the like.

[0156] The "means for transmitting collected conversations to a server" is a function for transmitting collected voice data to a server via a network.

[0157] The "means for analyzing the conversation on the server and generating appropriate phrases" is a function for analyzing the voice data received on the server and generating appropriate phrases based on the context and mood of the conversation.

[0158] The "means for transmitting the generated phrase to the terminal" is a function for transmitting the generated phrase to the terminal via the network.

[0159] The "means for reproducing a phrase in a whisper on the terminal" is a function that uses speech synthesis technology to reproduce the received phrase and whisper it to the user's ear.

[0160] "Means for analyzing content being viewed and providing commentary and supplementary information in real time" refers to a function that analyzes the video or audio content being viewed by the user and generates and provides commentary and supplementary information related to that content in real time.

[0161] The present invention provides a system that collects user conversations and provides appropriate phrases in real time. Specifically, it includes a function that also provides commentary and supplemental information related to the content being viewed in real time. The system of the present invention is implemented as follows.

[0162] 1. Audio collection and transmission

[0163] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user. This audio data is temporarily stored in the device's local memory and undergoes noise cancellation processing. The compressed audio data is then end-to-end encrypted and securely transmitted to a server using a low-latency communication protocol.

[0164] 2. Analysis and phrase generation on the server

[0165] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology. This textualized conversation undergoes contextual analysis and keyword extraction. It also simultaneously analyzes the user's mood, assessing the tone and emotional state of the conversation. Based on this data, appropriate phrases are generated using a generative AI model.

[0166] It also analyzes the content being viewed, analyzing the audio and video of the content and generating relevant commentary and supplemental information. For example, while watching educational content, it can provide academic background information and related new research results in real time.

[0167] 3. Sending and playing phrases

[0168] The generated phrase is then sent back to the device, which then converts the received phrase into voice using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear through the earphone speaker, allowing the user to obtain relevant information in real time while maintaining a natural conversation flow.

[0169] Specific examples

[0170] Situation 1: While watching an educational video

[0171] When a user is watching a scientific video and says, "Next, let's talk about Newtonian mechanics," the device collects this audio and sends it to the server. The server analyzes the content and generates the phrase, "Shall we also talk a bit about the law of gravity?" This allows the user to get additional information about the law of gravity in real time.

[0172] Situation 2: During a casual conversation

[0173] If a user says in a casual conversation, "What are you planning to do this weekend?", the server will generate a phrase like, "Why don't you invite her to try that new restaurant that everyone's talking about?" This allows users to naturally introduce new topics into their conversations.

[0174] Prompt Sentence Examples

[0175] (Example prompt):

[0176] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[0177] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[0178] The present invention enables users to receive appropriate phrases and information in real time that are in line with the content they are viewing or the context of the conversation, thereby realizing a rich communication environment.

[0179] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0180] Step 1:

[0181] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The input is the user's voice, and the output is clear audio data that has been processed with noise cancellation. This audio data is temporarily stored in the device's local memory.

[0182] Step 2:

[0183] The voice data stored on the device is compressed and sent to the server as data packets. The input is clear voice data, and the output is compressed data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted.

[0184] Step 3:

[0185] The server uses natural language processing (NLP) techniques to decode the received voice data and convert it into text. The input is compressed voice data, and the output is the text of the conversation. The server analyzes this text data and performs context analysis and keyword extraction.

[0186] Step 4:

[0187] The server uses a generative AI model to generate appropriate phrases based on the data obtained through context analysis and keyword extraction. The input is the data obtained through context analysis and keyword extraction, and the output is an appropriate phrase. This phrase also takes into account the user's mood.

[0188] Step 5:

[0189] The server sends the generated phrase to the terminal. The input is the generated phrase, and the output is the data sent to the terminal. The data is compressed and sent with the minimum amount of data. A low-latency communication protocol is used for transmission.

[0190] Step 6:

[0191] The device converts the received phrase into voice using Text-to-Speech (TTS) technology. The input is the phrase sent from the server, and the output is voice data. This voice data is played back as a "whisper" in the user's ear through the earphone speaker.

[0192] Step 7:

[0193] The content being viewed is also analyzed, and related commentary and supplementary information are generated. The input is the content data being viewed, and the output is commentary and supplementary information based on the analysis results. This data is generated in real time and provided to the user.

[0194] Example prompt

[0195] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[0196] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[0197] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0198] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then sent back to the terminal and "whispered" to suggest them to the user. The present invention also incorporates an emotion engine that recognizes the user's emotions, and generates appropriate phrases based on the emotional information.

[0199] 1. Recording conversations using a device

[0200] When a user wears the earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory and noise-canceling processing is applied. The audio data is organized into samples at regular time intervals and prepared for transmission to the server.

[0201] 2. Sending conversations by device

[0202] The device compresses the collected audio data and transmits it to the server as data packets. A low-latency communication protocol is used for transmission, maintaining real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[0203] 3. Analysis of conversations by the server

[0204] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques. The textual content undergoes contextual analysis and keyword extraction to identify important keywords and topics. The server also performs mood analysis to assess the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[0205] 4. Emotion Recognition by Emotion Engine

[0206] The server uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's voice characteristics, such as tone, pitch, rhythm, and speed, to determine multiple emotional states. This allows for a precise understanding of the user's current emotions.

[0207] 5. Server-generated phrases

[0208] The server generates appropriate and humorous phrases based on context and emotional information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. Based on a large amount of training data, this algorithm is capable of creating natural-sounding phrases that fit the context. The generated phrases are temporarily stored on the server and immediately sent to the device.

[0209] 6. Sending proposals from the server to the device

[0210] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[0211] 7. "Whisper" phrases from your device

[0212] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[0213] Specific examples

[0214] Example 1: Business meeting

[0215] If a user says, "I'd like to talk about a new project today," the device will pick up this speech and send it to the server. The server will extract the keywords "project" and "want to talk" and use an emotion engine to recognize emotions such as a mixture of tension and excitement. Based on this, the server will generate a phrase such as "It might be helpful if you also mention the progress of the project," and send it to the device. The phrase is suggested to the user in a "whispered" voice.

[0216] Example 2: Casual conversation

[0217] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the emotion of interest. The server generates the phrase "Why don't you try that cafe that's been getting a lot of attention lately?" and suggests it in a whisper from the device.

[0218] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations. By combining it with an emotion engine, more accurate suggestions that are in tune with the user's emotions become possible, improving the quality of communication.

[0219] The processing flow will be explained below.

[0220] Step 1:

[0221] The user wears the earphone-type device. The device uses a built-in microphone to pick up the conversations around the user in real time. The collected audio data is temporarily stored in the device's local memory and noise-canceling processing is applied.

[0222] Step 2:

[0223] The device samples the noise-canceled audio data at regular intervals, compresses it into data packets, and encrypts them. The encrypted audio data is then ready to be sent to the server.

[0224] Step 3:

[0225] The device sends encrypted data packets to the server using a low-latency communication protocol. To maintain real-time performance, the transmitted data is encrypted end-to-end.

[0226] Step 4:

[0227] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology, which is then temporarily stored on the server for analysis.

[0228] Step 5:

[0229] The server performs contextual analysis and keyword extraction based on the text data. Contextual analysis uses grammar rules and statistical models, while keyword extraction uses text mining techniques. After identifying important keywords and topics, the analysis results are passed on to the next processing step.

[0230] Step 6:

[0231] The server uses mood analysis and an emotion engine to recognize the user's emotional state. The emotion engine analyzes characteristics of the voice data, such as pitch, tone, rhythm, and speed, to evaluate the user's emotional state in real time. For example, multiple emotions such as joy, tension, and excitement can be recognized.

[0232] Step 7:

[0233] The server integrates contextual and emotional information to generate appropriate phrases. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. The algorithm is capable of creating natural-sounding phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and are then ready to be sent to the device.

[0234] Step 8:

[0235] The server compresses the generated phrase into a data packet and sends it to the device using a low-latency communication protocol. All data transmission is encrypted and occurs in real time.

[0236] Step 9:

[0237] The device decompresses the received phrase data and converts it into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker, allowing the user to use the suggested phrases in real time.

[0238] Step 10:

[0239] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data to the server.

[0240] Step 11:

[0241] The server then feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm, thus improving the system's ability to consistently deliver optimal phrases.

[0242] Example 2

[0243] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0244] In modern communication, it is important to provide appropriate conversations and responses instantly. However, users often find it difficult to think of appropriate phrases in real time. It is even more difficult to understand the conversation context and the user's emotions and make appropriate suggestions. Therefore, there is a need for a support system that helps users have smooth and engaging conversations.

[0245] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting voice data into text using natural language processing technology, and performing context analysis and keyword extraction, a means for performing emotion analysis, and a means for generating appropriate phrases based on the context and emotion information. This makes it possible to grasp the content of the user's conversation and emotional state in real time and provide appropriate and humorous phrases.

[0246] "User" refers to a person who uses the system to receive conversational assistance.

[0247] "Conversation" refers to verbal communication between a user and another person.

[0248] "Audio equipment" refers to a hardware device for collecting sound, such as an earphone-type device worn by a user.

[0249] "Sound collection" refers to capturing surrounding sounds in real time using audio equipment.

[0250] "Data compression" is the process of reducing the volume of audio data, which is done to increase transmission efficiency.

[0251] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[0252] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language.

[0253] "Text conversion" refers to the process of converting audio data into text data.

[0254] "Contextual analysis" refers to an analytical technique for understanding the content of a conversation and generating appropriate phrases based on that content.

[0255] "Keyword extraction" refers to the process of identifying and extracting important words and phrases from text.

[0256] "Emotion analysis" refers to techniques for identifying a user's emotional state from their vocal characteristics.

[0257] "Phrase generation" refers to the process of creating appropriate sentences based on the results of contextual and sentiment analysis.

[0258] "Data encryption technology" refers to technology that encrypts data to maintain its security.

[0259] A "low latency communication protocol" refers to a communication protocol that minimizes delays in sending and receiving data.

[0260] "Speech synthesis technology" refers to technology for converting text data into voice data.

[0261] "Whisper" refers to a soft voice generated by a terminal, and refers to a voice output format that sounds natural to the user.

[0262] The present invention is a system that uses an earphone-type device worn by the user to collect conversations in real time, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then transmitted back to the terminal and "whispered" to suggest them to the user. This system incorporates an emotion engine that recognizes the user's emotions, and has the advantage of generating appropriate phrases based on the emotion information. The specific steps for implementing the present invention are as follows.

[0263] The user first puts on the earphone-type device. The earphone has a built-in highly sensitive microphone that picks up surrounding conversations in real time. The audio data is temporarily stored in the device's local memory, and noise cancellation processing is performed at the same time. For example, a typical smartphone or portable audio device can be used as the device.

[0264] The device is equipped with data compression software, which compresses the collected audio data. The compressed data is then sent to the server using a low-latency communication protocol, such as WebRTC. The transmitted data is also end-to-end encrypted, protecting the user's privacy.

[0265] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques, possibly using the Google Cloud Speech-to-Text API or similar. The converted text undergoes contextual analysis and keyword extraction to extract important keywords and topics. The server then performs mood analysis to assess the tone and emotional state of the conversation. This analysis can be performed using the DeepAffects API or similar technologies.

[0266] Furthermore, the server is equipped with an emotion engine that recognizes multiple emotional states in real time by analyzing the tone, pitch, rhythm, and speed of the user's voice. Based on this information, it is possible to determine the emotion the user is feeling in the current conversation.

[0267] Based on contextual analysis and sentiment information, the server generates appropriate and humorous phrases using pre-trained generative AI models such as OpenAI's GPT-3, which are then temporarily stored on the server and immediately sent to the device.

[0268] The device converts the received phrase into voice using text-to-speech technology. Mimic3 or a similar voice synthesis engine can be used here. The generated voice data is played back as a "whisper" in the user's ear through the earphone speaker. This "whisper" voice allows the user to instantly use the appropriate phrase.

[0269] Through the above process, the present invention can provide users with appropriate phrases in real time, enabling smooth and engaging conversations in a variety of situations. Furthermore, by combining it with an emotion engine, more accurate suggestions become possible, improving the quality of communication.

[0270] Specific examples

[0271] Example 1: Business meeting

[0272] If a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts keywords like "project" and "want to talk" and uses an emotion engine to recognize tension and excitement. Based on this, the server generates a phrase like, "It might be helpful if you also mention the progress of the project," and sends it to the device. The user receives this suggestion in a "whispered" voice.

[0273] Example 2: Casual conversation

[0274] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the user's interests. The server generates the phrase "Why don't you try that cafe that's been trending lately?" and suggests it in a whisper from the device.

[0275] Prompt Sentence Examples

[0276] "Please suggest how users should talk about projects in business conversations."

[0277] "Show me appropriate phrases to use in casual conversation about weekend plans."

[0278] As a result, the present invention can provide users with appropriate phrases in real time to enable smooth and engaging conversations in a variety of situations.

[0279] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0280] Step 1:

[0281] After the user puts on the earphone-type device, the device activates the built-in microphone to collect surrounding sounds in real time. The input is the user's conversational voice. The device temporarily stores the collected voice data in local memory and performs noise cancellation processing. The output is the voice data after noise cancellation. Specifically, it uses the microphone function of a smartphone or portable audio device.

[0282] Step 2:

[0283] The device processes the noise-canceled audio data by arranging it as samples at regular time intervals and compressing it. The input is the noise-canceled audio data. For example, the G.711 compression format is used for data compression. The compressed data is sent to the server using a low-latency communication protocol (e.g., WebRTC). The output is compressed and encrypted audio data. Again, AES-256 technology is used for end-to-end encryption.

[0284] Step 3:

[0285] The server decodes the received compressed audio data. The input is compressed and encrypted audio data. After decoding, it is converted into PCM data. The server then converts the audio into text using natural language processing (NLP) technology (e.g., Google Cloud Speech-to-Text API). Specifically, the audio signal is analyzed and output as text data. The output is the text of the conversation.

[0286] Step 4:

[0287] The server performs contextual analysis and keyword extraction on the obtained text data. The input is the textual content of the conversation. The NLTK library and other libraries are used for contextual analysis and keyword extraction. As a result of the analysis, important keywords and topics are extracted. The output is keywords and contextual data.

[0288] Step 5:

[0289] The server then performs mood analysis. The input is text data and voice characteristics data. The DeepAffects API is used for emotion analysis. The tone, pitch, rhythm, and speed of the voice are analyzed to determine the user's emotional state. The output is data indicating the user's emotional state (e.g., "tense" or "excited").

[0290] Step 6:

[0291] The server generates appropriate phrases using a generative AI model (e.g., GPT-3) based on contextual analysis and emotional information. The inputs are keywords, contextual data, and emotional data. The generated phrases are temporarily stored on the server. The output is the generated phrase. Here, the behavior of the generative AI model is adjusted by using prompt sentences as examples.

[0292] Step 7:

[0293] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase. The OPUS format is used to compress the data. AES-256 encryption technology is also applied during transmission. The output is the compressed and encrypted phrase data.

[0294] Step 8:

[0295] The device decodes the received phrase and converts it into voice using speech synthesis technology (e.g., Mimic3). The input is compressed and encrypted phrase data. The speech synthesis generates a whispered voice. The generated voice is played to the user's ear through the earphone speaker. The output is a voiced phrase. This "whispered" voice allows the user to use appropriate phrases during actual conversations.

[0296] (Application example 2)

[0297] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0298] Modern autonomous vehicles lack a smooth and comfortable way to communicate with passengers. This can cause stress and anxiety for passengers, potentially reducing the quality of their riding experience. Furthermore, traditional voice assistants lack the ability to adequately analyze emotions and provide prompt and appropriate feedback. This can result in inappropriate responses to passenger instructions and requests.

[0299] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0300] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for playing the phrases in a whisper in the terminal, and means for generating appropriate feedback based on the collected conversation and the analysis results and suggesting the feedback to the user in a voice. This makes it possible to reduce the user's stress and anxiety while riding and provide a comfortable riding experience.

[0301] "User" refers to a person or end user who uses the system.

[0302] "Means for collecting conversation sounds" refers to equipment including a microphone and audio input device for collecting sounds around the user.

[0303] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[0304] "Means for transmitting the conversation to the server" refers to equipment that includes a communication module and protocol for transmitting collected voice data to a remote server over a network.

[0305] "Means for analyzing conversations" refers to natural language processing technologies and algorithms that convert collected audio data into text and analyze its content.

[0306] "Means for generating appropriate phrases" refers to generative AI models and algorithms that generate appropriate responses and suggestions based on analyzed text data and the user's emotional state.

[0307] "Means for transmitting phrases to a terminal" refers to a device that includes a communication module or protocol for transmitting the generated phrases to a user's terminal via a network.

[0308] "Means for playing in a whisper" refers to speakers or earphones that use voice synthesis technology to play phrases received by the terminal to the user at a low volume.

[0309] "Means for generating appropriate feedback and providing audible suggestions to the user" refers to a system in general for providing appropriate feedback and suggestions to the user audibly based on the analysis results.

[0310] To realize the present invention, the following system configuration and processing flow are included.

[0311] The present invention provides a method for collecting and analyzing a user's conversation in real time and suggesting appropriate feedback to the user by utilizing an earphone-type device, a microphone, a server, a communication protocol, and speech synthesis technology.

[0312] Hardware and Software Configuration

[0313] 1. User terminal (earphone-type device)

[0314] Microphone: Picks up the user's conversation and provides noise cancellation.

[0315] Communication module: A module for achieving low latency communication with the server.

[0316] Speaker / Earphone: Phrases received from the server are played back in a whisper using voice synthesis technology.

[0317] 2. Server

[0318] Natural language processing engine (NLP engine): Analyzes voice data and converts it into text.

[0319] Sentiment analysis engine: Analyzes user emotions from voice or text data.

[0320] Generative AI model: Generates appropriate feedback based on the user's conversational content and emotional state.

[0321] Communication module: A module for low-latency, encrypted communication with user devices.

[0322] System processing flow

[0323] 1. Audio collection:

[0324] When a user speaks, the microphone in the earphone-type device picks up the conversation, and the audio data is temporarily stored in the device's local memory and noise-canceling is performed.

[0325] 2. Sending audio data:

[0326] The collected audio data is compressed and encrypted before being sent to a server using a low-latency, highly secure protocol.

[0327] 3. Analysis of audio data:

[0328] The server receives the voice data and converts it into text using a natural language processing engine. It then analyzes the text data to extract important keywords and sentences.

[0329] 4. Emotion analysis:

[0330] In parallel with natural language processing, a sentiment analysis engine assesses the user's emotional state (e.g., happy, sad, nervous) in real time.

[0331] 5. Generate appropriate feedback:

[0332] Based on the analysis results and sentiment data, the generative AI model generates appropriate phrases that include helpful information and suggestions for the user.

[0333] 6. Sending the phrase from the server to the device:

[0334] The generated phrases are compressed and transmitted to the user terminal via a low-latency communication protocol.

[0335] 7. Phrase audio playback:

[0336] The user terminal uses voice synthesis technology to play back the received phrase in a "whispered" voice and suggest it to the user.

[0337] Specific examples

[0338] Example 1: Stress reduction function

[0339] If a passenger says, "I'm feeling a little stressed right now," the voice data is sent to the server. The server's emotion analysis engine detects "stress," and the generative AI model generates a suggestion: "Would you like me to play some music to change your mood?" The words "Would you like me to play some music to change your mood?" are whispered through the earphones.

[0340] Example 2: Tourist information function

[0341] When a passenger asks, "What are the tourist spots around here?", the server analyzes the keywords "tourism" and "spot." The generative AI model generates a suggestion such as, "There's a famous park nearby. Would you like me to show you around?", and a voice whispers through the earphones, "There's a famous park nearby. Would you like me to show you around?"

[0342] Prompt Sentence Examples

[0343] "Generate appropriate suggestions based on passenger statements. Example statement: 'I'm feeling a bit stressed right now.'"

[0344] "Generate suggestions based on the context and emotional state of the utterance. Example utterance: 'What are the tourist attractions around here?'"

[0345] In this way, the present invention provides a system that provides appropriate feedback to the user in real time, thereby realizing a comfortable riding experience for the user.

[0346] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0347] Step 1:

[0348] When a user speaks, the conversation is picked up by a microphone built into the terminal (earphone-type device). The collected audio data undergoes noise cancellation processing and is temporarily stored in local memory. The input is the user's conversational voice, and the output is noise-canceled audio data.

[0349] Step 2:

[0350] The device compresses the noise-canceled voice data and transmits it to the server using a low-latency, encrypted communication protocol. The input is the noise-canceled voice data, and the output is compressed and encrypted voice data packets.

[0351] Step 3:

[0352] The server receives the voice data packets sent from the device, decodes them, and converts them back into the original voice data. It then converts this voice data into text using a natural language processing engine. The input is the encrypted and compressed voice data packets, and the output is the text of the conversation.

[0353] Step 4:

[0354] The server analyzes the textual content of the conversation and extracts context and keywords. It then uses a sentiment analysis engine to evaluate the user's emotional state (e.g., joy, sadness, tension) from the text. The input is the textual content of the conversation, and the output is the extracted keywords and the user's emotional state.

[0355] Step 5:

[0356] The server uses a generative AI model to generate appropriate phrases based on the extracted keywords and emotional state. These generated phrases contain useful information and suggestions for the user. The input is keywords and emotional state, and the output is the generated appropriate phrases.

[0357] Step 6:

[0358] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase and the output is a compressed data packet.

[0359] Step 7:

[0360] The device decodes the data packets received from the server and converts the phrases into audio using text-to-speech technology. It then plays the suggestions to the user in a whispered voice through the earphone speaker. The input is the compressed data packets, and the output is the suggested audio played in a whispered voice.

[0361] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0362] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0363] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0364] [Second embodiment]

[0365] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0366] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0367] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0368] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0369] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0370] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0371] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0372] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0373] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0374] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0375] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0376] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0377] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, transmits the audio data to a server for analysis, and sends appropriate phrases generated by the analysis back to the terminal, where they are "whispered" to suggest the phrases to the user. The present invention is implemented as follows.

[0378] 1. Recording conversations using a device

[0379] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory. Noise cancellation processing is applied to obtain clear audio. The device prepares to send this audio data to a server at regular intervals.

[0380] 2. Sending conversations by device

[0381] The device compresses the voice data and sends it to the server as data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted to ensure security. Data transmission is always in real time, with great care taken to minimize latency.

[0382] 3. Analysis of conversations by the server

[0383] The server decodes the received audio data and converts it into text using natural language processing (NLP). The converted text undergoes contextual analysis and keyword extraction to identify important keywords and topics. Mood analysis is also performed to evaluate the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[0384] 4. Phrase generation by the server

[0385] The server generates appropriate and humorous phrases based on the extracted keywords and mood information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. This algorithm is capable of creating natural phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and immediately sent to the device.

[0386] 5. Proposal transmission from the server to the device

[0387] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[0388] 6. "Whisper" phrases from your device

[0389] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[0390] Specific examples

[0391] Example 1: Business meeting

[0392] When a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts the keywords "project" and "want to talk," performs a mood analysis, and generates an appropriate phrase, such as "Why don't you talk a little about the progress of the project?" The generated phrase is sent to the device and played as a "whisper" in the user's ear.

[0393] Example 2: Casual conversation

[0394] When a user says, "Do you have any plans for the weekend?", the device collects the corresponding voice and sends it to the server. The server analyzes "weekend" and "plans" as keywords and also considers the user's emotional state (e.g., interests), generating the phrase "Why don't you invite me to a new restaurant that's been making waves lately?". This phrase is then whispered through the device: "Do you want to try a restaurant that's been making waves lately?"

[0395] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations.

[0396] The processing flow will be explained below.

[0397] Step 1:

[0398] The user wears the earphone-type device. The device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory.

[0399] Step 2:

[0400] The device applies noise cancellation to the collected voice data to make it clear. The voice data is sampled at regular intervals, compressed into data packets, and encrypted.

[0401] Step 3:

[0402] The device sends compressed and encrypted audio data to the server using a low-latency communication protocol to maintain real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[0403] Step 4:

[0404] The server decodes the received voice data and converts it into text using natural language processing (NLP) techniques.

[0405] Step 5:

[0406] The server performs contextual analysis and keyword extraction on the converted conversation content, identifying important keywords and topics and understanding the context of the conversation.

[0407] Step 6:

[0408] The server performs mood analysis to assess the tone and emotional state of the conversation, thereby capturing the atmosphere of the conversation (e.g., joy, tension, excitement).

[0409] Step 7:

[0410] The server generates appropriate and humorous phrases based on context and mood information, using pre-trained artificial intelligence (AI) algorithms.

[0411] Step 8:

[0412] The server then re-compresses the generated phrases into data packets, which are then sent to the device using a low-latency communication protocol.

[0413] Step 9:

[0414] The device decompresses the received phrase data and converts it from text to voice using text-to-speech technology.

[0415] Step 10:

[0416] The device then plays the generated voice data through the earphone speaker in a whisper close to the user's ear, allowing the user to continue the conversation using phrases suggested in real time.

[0417] Step 11:

[0418] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data back to the server.

[0419] Step 12:

[0420] The server feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm.

[0421] Example 1

[0422] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0423] In many modern conversational environments, users face the challenge of generating appropriate phrases in real time and smoothly advancing a conversation. In particular, in business meetings and casual conversations, being able to instantly respond appropriately to the situation requires high communication skills. Therefore, there is a need for technology that allows users to obtain appropriate phrases in a timely manner during a conversation.

[0424] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0425] In this invention, the server includes a means for analyzing speech and converting it into text data, a means for using a generative AI model to generate appropriate phrases based on the text data, and a means for analyzing the user's emotional state, thereby enabling the server to analyze the user's conversation in real time and instantly generate and provide appropriate and humorous phrases.

[0426] "User" refers to an individual who wears an earphone-type device and is the conversation partner.

[0427] "Voice" refers to conversations and other vocal content spoken by a user.

[0428] "Collection means" refers to the equipment and technology used to pick up and appropriately process audio.

[0429] A "terminal" is an earphone-type device worn by a user, and refers to a device that collects, processes, and transmits audio.

[0430] "Server" refers to a computer system for receiving voice data, analyzing it, converting it into text, generating phrases, etc.

[0431] "Generative AI model" refers to an algorithm equipped with artificial intelligence technology that is used to generate appropriate phrases based on text data.

[0432] "Audio playback means" refers to a technique or device that plays back the generated phrase as audio.

[0433] "Natural language processing technology" refers to computer technology for processing, understanding, and generating human language.

[0434] "Emotional state" refers to the user's psychological state, which indicates the tone and emotional state of the conversation.

[0435] The present invention relates to a system in which a user wears an earphone-type device, collects the sound of a conversation in real time, and transmits the collected sound data to a server for analysis. The system of the present invention is implemented as follows.

[0436] 1. The user puts on the device

[0437] When a user puts the earphone-type device in their ear, the device automatically starts up and turns on sound collection mode, allowing it to collect conversations around the user in real time.

[0438] 2. Recording conversations using a device

[0439] The device is equipped with a highly sensitive microphone that collects conversations around the user in real time. The audio data is temporarily stored in the device's local memory, and noise-cancelling technology is used to ensure clear audio. This process uses a microphone with noise-cancelling technology.

[0440] 3. Transmission of audio data by the device

[0441] The device compresses the collected audio data and sends it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end to ensure security. An audio compression codec such as G.711 is used for this compression.

[0442] 4. Receiving and analyzing audio data by the server

[0443] The server receives the voice data sent from the device and uses a speech recognition engine (e.g., Google Speech-to-Text) to decode it. The server converts the voice data into text data, then uses natural language processing (NLP) techniques to analyze the context of the voice and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state.

[0444] 5. Server-generated phrases

[0445] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. This generative AI model has the ability to predict and generate the most appropriate phrases based on the user's speech content and emotional state.

[0446] 6. Sending proposals from the server to the device

[0447] The generated phrase is then compressed again and sent to the device using a low-latency communication protocol. The device then decodes the compressed data and moves on to the next step.

[0448] 7. "Whisper" phrases from your device

[0449] The device converts the received phrase into a voice in real time using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear. This voice playback uses a TTS engine (e.g., Amazon Polly).

[0450] Specific examples

[0451] Example 1: Business meeting

[0452] User: "Today I'd like to talk about a new project."

[0453] The device collects the audio and sends it to the server.

[0454] The server analyzes keywords such as "project" and "want to talk" and performs mood analysis.

[0455] The server uses a "generative AI model" to generate the phrase "Why not talk a little about the progress of the project?" and sends it to the device.

[0456] The terminal whispers, "Why not give us a little update on the progress of the project?"

[0457] Example 2: Casual conversation

[0458] User: "Do you have any plans for the weekend?"

[0459] The device collects the audio and sends it to the server.

[0460] The server analyzes "weekend" and "schedule" and takes into account the user's emotional state.

[0461] The server uses a "generative AI model" to generate the phrase "Why not invite me to a new restaurant that's the talk of the town?" and sends it to the device.

[0462] The device whispers, "Would you like to try a restaurant that's been getting a lot of attention lately?"

[0463] In this way, the present invention provides users with appropriate and humorous phrases in real time, allowing for smooth and engaging conversations.

[0464] Example prompt sentence:

[0465] "When a user says, 'Today I'd like to talk about a new project' in a business meeting, please suggest an appropriate phrase."

[0466] Example output:

[0467] "Maybe we should talk a little bit about how the project is going."

[0468] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0469] Step 1:

[0470] User Wears Device

[0471] The user puts the earphone-type device in their ear. This causes the device to automatically start up and turn on sound collection mode. The input is the state in which the earphone is being worn, and the output is the device switching to sound collection mode.

[0472] Step 2:

[0473] Recording conversations using a device

[0474] The device uses a built-in microphone to collect conversations around the user in real time. The collected audio data is temporarily stored in local memory, and clear audio is obtained using noise-canceling technology. The input is the surrounding audio, and the output is clear audio data that has been subjected to noise-canceling processing. Specifically, the device collects audio every second and stores the data in a buffer.

[0475] Step 3:

[0476] Sending audio data by the device

[0477] The device compresses the collected voice data using a voice compression codec such as G.711 and transmits it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end before transmission. The input is clear voice data, and the output is compressed and encrypted data packets.

[0478] Step 4:

[0479] Receiving and analyzing voice data by the server

[0480] The server receives and decodes the voice data sent from the device. It then uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. It then uses natural language processing (NLP) technology to analyze the context of the text and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state. The input is the decoded voice data, and the output is the analyzed text and emotional information. Specifically, the server receives the voice data, converts it into text through the speech recognition engine, and then analyzes it using an NLP model.

[0481] Step 5:

[0482] Server-generated phrases

[0483] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. The generative AI model predicts and generates the most appropriate phrase based on the prompt sentence. The input is the text data and emotional information, and the output is the generated phrase. Specifically, the server inputs the prompt sentence into the generative AI model and obtains the generated phrase.

[0484] Step 6:

[0485] Sending proposals from the server to the device

[0486] The server compresses the generated phrase and transmits it to the terminal using a low-latency communication protocol. The input is the generated phrase, and the output is a compressed and encrypted data packet. Specifically, the server compresses the generated phrase and encrypts it for transmission as a data packet.

[0487] Step 7:

[0488] "Whisper" phrases from your device

[0489] The device converts the received phrase into audio in real time using Text-to-Speech (TTS) technology and plays it as a "whisper" in the user's ear. The input is encrypted phrase data, and the output is a spoken suggested phrase. Specifically, the device uses a TTS engine to convert the phrase into audio and plays the audio through earphones.

[0490] (Application example 1)

[0491] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0492] In conventional content distribution services, it has been difficult to provide appropriate commentary or supplemental information about the content being viewed in real time. It has also been a challenge to provide accurate information according to the user's conversation and mood. The present invention aims to solve these problems.

[0493] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0494] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for reproducing the phrases in a whisper in the terminal, and means for analyzing the content being viewed and providing commentary and supplemental information in real time, thereby enabling the user to enjoy appropriate commentary and supplemental information in real time even while viewing the content.

[0495] The "means for collecting the user's conversation" is a function for collecting the user's voice using a microphone or the like.

[0496] The "means for transmitting collected conversations to a server" is a function for transmitting collected voice data to a server via a network.

[0497] The "means for analyzing the conversation on the server and generating appropriate phrases" is a function for analyzing the voice data received on the server and generating appropriate phrases based on the context and mood of the conversation.

[0498] The "means for transmitting the generated phrase to the terminal" is a function for transmitting the generated phrase to the terminal via the network.

[0499] The "means for reproducing a phrase in a whisper on the terminal" is a function that uses speech synthesis technology to reproduce the received phrase and whisper it to the user's ear.

[0500] "Means for analyzing content being viewed and providing commentary and supplementary information in real time" refers to a function that analyzes the video or audio content being viewed by the user and generates and provides commentary and supplementary information related to that content in real time.

[0501] The present invention provides a system that collects user conversations and provides appropriate phrases in real time. Specifically, it includes a function that also provides commentary and supplemental information related to the content being viewed in real time. The system of the present invention is implemented as follows.

[0502] 1. Audio collection and transmission

[0503] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user. This audio data is temporarily stored in the device's local memory and undergoes noise cancellation processing. The compressed audio data is then end-to-end encrypted and securely transmitted to a server using a low-latency communication protocol.

[0504] 2. Analysis and phrase generation on the server

[0505] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology. This textualized conversation undergoes contextual analysis and keyword extraction. It also simultaneously analyzes the user's mood, assessing the tone and emotional state of the conversation. Based on this data, appropriate phrases are generated using a generative AI model.

[0506] It also analyzes the content being viewed, analyzing the audio and video of the content and generating relevant commentary and supplemental information. For example, while watching educational content, it can provide academic background information and related new research results in real time.

[0507] 3. Sending and playing phrases

[0508] The generated phrase is then sent back to the device, which then converts the received phrase into voice using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear through the earphone speaker, allowing the user to obtain relevant information in real time while maintaining a natural conversation flow.

[0509] Specific examples

[0510] Situation 1: While watching an educational video

[0511] When a user is watching a scientific video and says, "Next, let's talk about Newtonian mechanics," the device collects this audio and sends it to the server. The server analyzes the content and generates the phrase, "Shall we also talk a bit about the law of gravity?" This allows the user to get additional information about the law of gravity in real time.

[0512] Situation 2: During a casual conversation

[0513] If a user says in a casual conversation, "What are you planning to do this weekend?", the server will generate a phrase like, "Why don't you invite her to try that new restaurant that everyone's talking about?" This allows users to naturally introduce new topics into their conversations.

[0514] Prompt Sentence Examples

[0515] (Example prompt):

[0516] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[0517] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[0518] The present invention enables users to receive appropriate phrases and information in real time that are in line with the content they are viewing or the context of the conversation, thereby realizing a rich communication environment.

[0519] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0520] Step 1:

[0521] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The input is the user's voice, and the output is clear audio data that has been processed with noise cancellation. This audio data is temporarily stored in the device's local memory.

[0522] Step 2:

[0523] The voice data stored on the device is compressed and sent to the server as data packets. The input is clear voice data, and the output is compressed data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted.

[0524] Step 3:

[0525] The server uses natural language processing (NLP) techniques to decode the received voice data and convert it into text. The input is compressed voice data, and the output is the text of the conversation. The server analyzes this text data and performs context analysis and keyword extraction.

[0526] Step 4:

[0527] The server uses a generative AI model to generate appropriate phrases based on the data obtained through context analysis and keyword extraction. The input is the data obtained through context analysis and keyword extraction, and the output is an appropriate phrase. This phrase also takes into account the user's mood.

[0528] Step 5:

[0529] The server sends the generated phrase to the terminal. The input is the generated phrase, and the output is the data sent to the terminal. The data is compressed and sent with the minimum amount of data. A low-latency communication protocol is used for transmission.

[0530] Step 6:

[0531] The device converts the received phrase into voice using Text-to-Speech (TTS) technology. The input is the phrase sent from the server, and the output is voice data. This voice data is played back as a "whisper" in the user's ear through the earphone speaker.

[0532] Step 7:

[0533] The content being viewed is also analyzed, and related commentary and supplementary information are generated. The input is the content data being viewed, and the output is commentary and supplementary information based on the analysis results. This data is generated in real time and provided to the user.

[0534] Example prompt

[0535] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[0536] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[0537] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0538] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then sent back to the terminal and "whispered" to suggest them to the user. The present invention also incorporates an emotion engine that recognizes the user's emotions, and generates appropriate phrases based on the emotional information.

[0539] 1. Recording conversations using a device

[0540] When a user wears the earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory and noise-canceling processing is applied. The audio data is organized into samples at regular time intervals and prepared for transmission to the server.

[0541] 2. Sending conversations by device

[0542] The device compresses the collected audio data and transmits it to the server as data packets. A low-latency communication protocol is used for transmission, maintaining real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[0543] 3. Analysis of conversations by the server

[0544] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques. The textual content undergoes contextual analysis and keyword extraction to identify important keywords and topics. The server also performs mood analysis to assess the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[0545] 4. Emotion Recognition by Emotion Engine

[0546] The server uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's voice characteristics, such as tone, pitch, rhythm, and speed, to determine multiple emotional states. This allows for a precise understanding of the user's current emotions.

[0547] 5. Server-generated phrases

[0548] The server generates appropriate and humorous phrases based on context and emotional information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. Based on a large amount of training data, this algorithm is capable of creating natural-sounding phrases that fit the context. The generated phrases are temporarily stored on the server and immediately sent to the device.

[0549] 6. Sending proposals from the server to the device

[0550] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[0551] 7. "Whisper" phrases from your device

[0552] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[0553] Specific examples

[0554] Example 1: Business meeting

[0555] If a user says, "I'd like to talk about a new project today," the device will pick up this speech and send it to the server. The server will extract the keywords "project" and "want to talk" and use an emotion engine to recognize emotions such as a mixture of tension and excitement. Based on this, the server will generate a phrase such as "It might be helpful if you also mention the progress of the project," and send it to the device. The phrase is suggested to the user in a "whispered" voice.

[0556] Example 2: Casual conversation

[0557] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the emotion of interest. The server generates the phrase "Why don't you try that cafe that's been getting a lot of attention lately?" and suggests it in a whisper from the device.

[0558] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations. By combining it with an emotion engine, more accurate suggestions that are in tune with the user's emotions become possible, improving the quality of communication.

[0559] The processing flow will be explained below.

[0560] Step 1:

[0561] The user wears the earphone-type device. The device uses a built-in microphone to pick up the conversations around the user in real time. The collected audio data is temporarily stored in the device's local memory and noise-canceling processing is applied.

[0562] Step 2:

[0563] The device samples the noise-canceled audio data at regular intervals, compresses it into data packets, and encrypts them. The encrypted audio data is then ready to be sent to the server.

[0564] Step 3:

[0565] The device sends encrypted data packets to the server using a low-latency communication protocol. To maintain real-time performance, the transmitted data is encrypted end-to-end.

[0566] Step 4:

[0567] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology, which is then temporarily stored on the server for analysis.

[0568] Step 5:

[0569] The server performs contextual analysis and keyword extraction based on the text data. Contextual analysis uses grammar rules and statistical models, while keyword extraction uses text mining techniques. After identifying important keywords and topics, the analysis results are passed on to the next processing step.

[0570] Step 6:

[0571] The server uses mood analysis and an emotion engine to recognize the user's emotional state. The emotion engine analyzes characteristics of the voice data, such as pitch, tone, rhythm, and speed, to evaluate the user's emotional state in real time. For example, multiple emotions such as joy, tension, and excitement can be recognized.

[0572] Step 7:

[0573] The server integrates contextual and emotional information to generate appropriate phrases. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. The algorithm is capable of creating natural-sounding phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and are then ready to be sent to the device.

[0574] Step 8:

[0575] The server compresses the generated phrase into a data packet and sends it to the device using a low-latency communication protocol. All data transmission is encrypted and occurs in real time.

[0576] Step 9:

[0577] The device decompresses the received phrase data and converts it into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker, allowing the user to use the suggested phrases in real time.

[0578] Step 10:

[0579] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data to the server.

[0580] Step 11:

[0581] The server then feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm, thus improving the system's ability to consistently deliver optimal phrases.

[0582] Example 2

[0583] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0584] In modern communication, it is important to provide appropriate conversations and responses instantly. However, users often find it difficult to think of appropriate phrases in real time. It is even more difficult to understand the conversation context and the user's emotions and make appropriate suggestions. Therefore, there is a need for a support system that helps users have smooth and engaging conversations.

[0585] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting voice data into text using natural language processing technology, and performing context analysis and keyword extraction, a means for performing emotion analysis, and a means for generating appropriate phrases based on the context and emotion information. This makes it possible to grasp the content of the user's conversation and emotional state in real time and provide appropriate and humorous phrases.

[0586] "User" refers to a person who uses the system to receive conversational assistance.

[0587] "Conversation" refers to verbal communication between a user and another person.

[0588] "Audio equipment" refers to a hardware device for collecting sound, such as an earphone-type device worn by a user.

[0589] "Sound collection" refers to capturing surrounding sounds in real time using audio equipment.

[0590] "Data compression" is the process of reducing the volume of audio data, which is done to increase transmission efficiency.

[0591] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[0592] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language.

[0593] "Text conversion" refers to the process of converting audio data into text data.

[0594] "Contextual analysis" refers to an analytical technique for understanding the content of a conversation and generating appropriate phrases based on that content.

[0595] "Keyword extraction" refers to the process of identifying and extracting important words and phrases from text.

[0596] "Emotion analysis" refers to techniques for identifying a user's emotional state from their vocal characteristics.

[0597] "Phrase generation" refers to the process of creating appropriate sentences based on the results of contextual and sentiment analysis.

[0598] "Data encryption technology" refers to technology that encrypts data to maintain its security.

[0599] A "low latency communication protocol" refers to a communication protocol that minimizes delays in sending and receiving data.

[0600] "Speech synthesis technology" refers to technology for converting text data into voice data.

[0601] "Whisper" refers to a soft voice generated by a terminal, and refers to a voice output format that sounds natural to the user.

[0602] The present invention is a system that uses an earphone-type device worn by the user to collect conversations in real time, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then transmitted back to the terminal and "whispered" to suggest them to the user. This system incorporates an emotion engine that recognizes the user's emotions, and has the advantage of generating appropriate phrases based on the emotion information. The specific steps for implementing the present invention are as follows.

[0603] The user first puts on the earphone-type device. The earphone has a built-in highly sensitive microphone that picks up surrounding conversations in real time. The audio data is temporarily stored in the device's local memory, and noise cancellation processing is performed at the same time. For example, a typical smartphone or portable audio device can be used as the device.

[0604] The device is equipped with data compression software, which compresses the collected audio data. The compressed data is then sent to the server using a low-latency communication protocol, such as WebRTC. The transmitted data is also end-to-end encrypted, protecting the user's privacy.

[0605] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques, possibly using the Google Cloud Speech-to-Text API or similar. The converted text undergoes contextual analysis and keyword extraction to extract important keywords and topics. The server then performs mood analysis to assess the tone and emotional state of the conversation. This analysis can be performed using the DeepAffects API or similar technologies.

[0606] Furthermore, the server is equipped with an emotion engine that recognizes multiple emotional states in real time by analyzing the tone, pitch, rhythm, and speed of the user's voice. Based on this information, it is possible to determine the emotion the user is feeling in the current conversation.

[0607] Based on contextual analysis and sentiment information, the server generates appropriate and humorous phrases using pre-trained generative AI models such as OpenAI's GPT-3, which are then temporarily stored on the server and immediately sent to the device.

[0608] The device converts the received phrase into voice using text-to-speech technology. Mimic3 or a similar voice synthesis engine can be used here. The generated voice data is played back as a "whisper" in the user's ear through the earphone speaker. This "whisper" voice allows the user to instantly use the appropriate phrase.

[0609] Through the above process, the present invention can provide users with appropriate phrases in real time, enabling smooth and engaging conversations in a variety of situations. Furthermore, by combining it with an emotion engine, more accurate suggestions become possible, improving the quality of communication.

[0610] Specific examples

[0611] Example 1: Business meeting

[0612] If a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts keywords like "project" and "want to talk" and uses an emotion engine to recognize tension and excitement. Based on this, the server generates a phrase like, "It might be helpful if you also mention the progress of the project," and sends it to the device. The user receives this suggestion in a "whispered" voice.

[0613] Example 2: Casual conversation

[0614] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the user's interests. The server generates the phrase "Why don't you try that cafe that's been trending lately?" and suggests it in a whisper from the device.

[0615] Prompt Sentence Examples

[0616] "Please suggest how users should talk about projects in business conversations."

[0617] "Show me appropriate phrases to use in casual conversation about weekend plans."

[0618] As a result, the present invention can provide users with appropriate phrases in real time to enable smooth and engaging conversations in a variety of situations.

[0619] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0620] Step 1:

[0621] After the user puts on the earphone-type device, the device activates the built-in microphone to collect surrounding sounds in real time. The input is the user's conversational voice. The device temporarily stores the collected voice data in local memory and performs noise cancellation processing. The output is the voice data after noise cancellation. Specifically, it uses the microphone function of a smartphone or portable audio device.

[0622] Step 2:

[0623] The device processes the noise-canceled audio data by arranging it as samples at regular time intervals and compressing it. The input is the noise-canceled audio data. For example, the G.711 compression format is used for data compression. The compressed data is sent to the server using a low-latency communication protocol (e.g., WebRTC). The output is compressed and encrypted audio data. Again, AES-256 technology is used for end-to-end encryption.

[0624] Step 3:

[0625] The server decodes the received compressed audio data. The input is compressed and encrypted audio data. After decoding, it is converted into PCM data. The server then converts the audio into text using natural language processing (NLP) technology (e.g., Google Cloud Speech-to-Text API). Specifically, the audio signal is analyzed and output as text data. The output is the text of the conversation.

[0626] Step 4:

[0627] The server performs contextual analysis and keyword extraction on the obtained text data. The input is the textual content of the conversation. The NLTK library and other libraries are used for contextual analysis and keyword extraction. As a result of the analysis, important keywords and topics are extracted. The output is keywords and contextual data.

[0628] Step 5:

[0629] The server then performs mood analysis. The input is text data and voice characteristics data. The DeepAffects API is used for emotion analysis. The tone, pitch, rhythm, and speed of the voice are analyzed to determine the user's emotional state. The output is data indicating the user's emotional state (e.g., "tense" or "excited").

[0630] Step 6:

[0631] The server generates appropriate phrases using a generative AI model (e.g., GPT-3) based on contextual analysis and emotional information. The inputs are keywords, contextual data, and emotional data. The generated phrases are temporarily stored on the server. The output is the generated phrase. Here, the behavior of the generative AI model is adjusted by using prompt sentences as examples.

[0632] Step 7:

[0633] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase. The OPUS format is used to compress the data. AES-256 encryption technology is also applied during transmission. The output is the compressed and encrypted phrase data.

[0634] Step 8:

[0635] The device decodes the received phrase and converts it into voice using speech synthesis technology (e.g., Mimic3). The input is compressed and encrypted phrase data. The speech synthesis generates a whispered voice. The generated voice is played to the user's ear through the earphone speaker. The output is a voiced phrase. This "whispered" voice allows the user to use appropriate phrases during actual conversations.

[0636] (Application example 2)

[0637] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0638] Modern autonomous vehicles lack a smooth and comfortable way to communicate with passengers. This can cause stress and anxiety for passengers, potentially reducing the quality of their riding experience. Furthermore, traditional voice assistants lack the ability to adequately analyze emotions and provide prompt and appropriate feedback. This can result in inappropriate responses to passenger instructions and requests.

[0639] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0640] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for playing the phrases in a whisper in the terminal, and means for generating appropriate feedback based on the collected conversation and the analysis results and suggesting the feedback to the user in a voice. This makes it possible to reduce the user's stress and anxiety while riding and provide a comfortable riding experience.

[0641] "User" refers to a person or end user who uses the system.

[0642] "Means for collecting conversation sounds" refers to equipment including a microphone and audio input device for collecting sounds around the user.

[0643] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[0644] "Means for transmitting the conversation to the server" refers to equipment that includes a communication module and protocol for transmitting collected voice data to a remote server over a network.

[0645] "Means for analyzing conversations" refers to natural language processing technologies and algorithms that convert collected audio data into text and analyze its content.

[0646] "Means for generating appropriate phrases" refers to generative AI models and algorithms that generate appropriate responses and suggestions based on analyzed text data and the user's emotional state.

[0647] "Means for transmitting phrases to a terminal" refers to a device that includes a communication module or protocol for transmitting the generated phrases to a user's terminal via a network.

[0648] "Means for playing in a whisper" refers to speakers or earphones that use voice synthesis technology to play phrases received by the terminal to the user at a low volume.

[0649] "Means for generating appropriate feedback and providing audible suggestions to the user" refers to a system in general for providing appropriate feedback and suggestions to the user audibly based on the analysis results.

[0650] To realize the present invention, the following system configuration and processing flow are included.

[0651] The present invention provides a method for collecting and analyzing a user's conversation in real time and suggesting appropriate feedback to the user by utilizing an earphone-type device, a microphone, a server, a communication protocol, and speech synthesis technology.

[0652] Hardware and Software Configuration

[0653] 1. User terminal (earphone-type device)

[0654] Microphone: Picks up the user's conversation and provides noise cancellation.

[0655] Communication module: A module for achieving low latency communication with the server.

[0656] Speaker / Earphone: Phrases received from the server are played back in a whisper using voice synthesis technology.

[0657] 2. Server

[0658] Natural language processing engine (NLP engine): Analyzes voice data and converts it into text.

[0659] Sentiment analysis engine: Analyzes user emotions from voice or text data.

[0660] Generative AI model: Generates appropriate feedback based on the user's conversational content and emotional state.

[0661] Communication module: A module for low-latency, encrypted communication with user devices.

[0662] System processing flow

[0663] 1. Audio collection:

[0664] When a user speaks, the microphone in the earphone-type device picks up the conversation, and the audio data is temporarily stored in the device's local memory and noise-canceling is performed.

[0665] 2. Sending audio data:

[0666] The collected audio data is compressed and encrypted before being sent to a server using a low-latency, highly secure protocol.

[0667] 3. Analysis of audio data:

[0668] The server receives the voice data and converts it into text using a natural language processing engine. It then analyzes the text data to extract important keywords and sentences.

[0669] 4. Emotion analysis:

[0670] In parallel with natural language processing, a sentiment analysis engine assesses the user's emotional state (e.g., happy, sad, nervous) in real time.

[0671] 5. Generate appropriate feedback:

[0672] Based on the analysis results and sentiment data, the generative AI model generates appropriate phrases that include helpful information and suggestions for the user.

[0673] 6. Sending the phrase from the server to the device:

[0674] The generated phrases are compressed and transmitted to the user terminal via a low-latency communication protocol.

[0675] 7. Phrase audio playback:

[0676] The user terminal uses voice synthesis technology to play back the received phrase in a "whispered" voice and suggest it to the user.

[0677] Specific examples

[0678] Example 1: Stress reduction function

[0679] If a passenger says, "I'm feeling a little stressed right now," the voice data is sent to the server. The server's emotion analysis engine detects "stress," and the generative AI model generates a suggestion: "Would you like me to play some music to change your mood?" The words "Would you like me to play some music to change your mood?" are whispered through the earphones.

[0680] Example 2: Tourist information function

[0681] When a passenger asks, "What are the tourist spots around here?", the server analyzes the keywords "tourism" and "spot." The generative AI model generates a suggestion such as, "There's a famous park nearby. Would you like me to show you around?", and a voice whispers through the earphones, "There's a famous park nearby. Would you like me to show you around?"

[0682] Prompt Sentence Examples

[0683] "Generate appropriate suggestions based on passenger statements. Example statement: 'I'm feeling a bit stressed right now.'"

[0684] "Generate suggestions based on the context and emotional state of the utterance. Example utterance: 'What are the tourist attractions around here?'"

[0685] In this way, the present invention provides a system that provides appropriate feedback to the user in real time, thereby realizing a comfortable riding experience for the user.

[0686] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0687] Step 1:

[0688] When a user speaks, the conversation is picked up by a microphone built into the terminal (earphone-type device). The collected audio data undergoes noise cancellation processing and is temporarily stored in local memory. The input is the user's conversational voice, and the output is noise-canceled audio data.

[0689] Step 2:

[0690] The device compresses the noise-canceled voice data and transmits it to the server using a low-latency, encrypted communication protocol. The input is the noise-canceled voice data, and the output is compressed and encrypted voice data packets.

[0691] Step 3:

[0692] The server receives the voice data packets sent from the device, decodes them, and converts them back into the original voice data. It then converts this voice data into text using a natural language processing engine. The input is the encrypted and compressed voice data packets, and the output is the text of the conversation.

[0693] Step 4:

[0694] The server analyzes the textual content of the conversation and extracts context and keywords. It then uses a sentiment analysis engine to evaluate the user's emotional state (e.g., joy, sadness, tension) from the text. The input is the textual content of the conversation, and the output is the extracted keywords and the user's emotional state.

[0695] Step 5:

[0696] The server uses a generative AI model to generate appropriate phrases based on the extracted keywords and emotional state. These generated phrases contain useful information and suggestions for the user. The input is keywords and emotional state, and the output is the generated appropriate phrases.

[0697] Step 6:

[0698] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase and the output is a compressed data packet.

[0699] Step 7:

[0700] The device decodes the data packets received from the server and converts the phrases into audio using text-to-speech technology. It then plays the suggestions to the user in a whispered voice through the earphone speaker. The input is the compressed data packets, and the output is the suggested audio played in a whispered voice.

[0701] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0702] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0703] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0704] [Third embodiment]

[0705] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0706] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0707] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0708] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0709] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0710] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0711] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0712] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0713] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0714] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0715] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0716] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0717] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, transmits the audio data to a server for analysis, and sends appropriate phrases generated by the analysis back to the terminal, where they are "whispered" to suggest the phrases to the user. The present invention is implemented as follows.

[0718] 1. Recording conversations using a device

[0719] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory. Noise cancellation processing is applied to obtain clear audio. The device prepares to send this audio data to a server at regular intervals.

[0720] 2. Sending conversations by device

[0721] The device compresses the voice data and sends it to the server as data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted to ensure security. Data transmission is always in real time, with great care taken to minimize latency.

[0722] 3. Analysis of conversations by the server

[0723] The server decodes the received audio data and converts it into text using natural language processing (NLP). The converted text undergoes contextual analysis and keyword extraction to identify important keywords and topics. Mood analysis is also performed to evaluate the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[0724] 4. Phrase generation by the server

[0725] The server generates appropriate and humorous phrases based on the extracted keywords and mood information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. This algorithm is capable of creating natural phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and immediately sent to the device.

[0726] 5. Proposal transmission from the server to the device

[0727] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[0728] 6. "Whisper" phrases from your device

[0729] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[0730] Specific examples

[0731] Example 1: Business meeting

[0732] When a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts the keywords "project" and "want to talk," performs a mood analysis, and generates an appropriate phrase, such as "Why don't you talk a little about the progress of the project?" The generated phrase is sent to the device and played as a "whisper" in the user's ear.

[0733] Example 2: Casual conversation

[0734] When a user says, "Do you have any plans for the weekend?", the device collects the corresponding voice and sends it to the server. The server analyzes "weekend" and "plans" as keywords and also considers the user's emotional state (e.g., interests), generating the phrase "Why don't you invite me to a new restaurant that's been making waves lately?". This phrase is then whispered through the device: "Do you want to try a restaurant that's been making waves lately?"

[0735] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations.

[0736] The processing flow will be explained below.

[0737] Step 1:

[0738] The user wears the earphone-type device. The device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory.

[0739] Step 2:

[0740] The device applies noise cancellation to the collected voice data to make it clear. The voice data is sampled at regular intervals, compressed into data packets, and encrypted.

[0741] Step 3:

[0742] The device sends compressed and encrypted audio data to the server using a low-latency communication protocol to maintain real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[0743] Step 4:

[0744] The server decodes the received voice data and converts it into text using natural language processing (NLP) techniques.

[0745] Step 5:

[0746] The server performs contextual analysis and keyword extraction on the converted conversation content, identifying important keywords and topics and understanding the context of the conversation.

[0747] Step 6:

[0748] The server performs mood analysis to assess the tone and emotional state of the conversation, thereby capturing the atmosphere of the conversation (e.g., joy, tension, excitement).

[0749] Step 7:

[0750] The server generates appropriate and humorous phrases based on context and mood information, using pre-trained artificial intelligence (AI) algorithms.

[0751] Step 8:

[0752] The server then re-compresses the generated phrases into data packets, which are then sent to the device using a low-latency communication protocol.

[0753] Step 9:

[0754] The device decompresses the received phrase data and converts it from text to voice using text-to-speech technology.

[0755] Step 10:

[0756] The device then plays the generated voice data through the earphone speaker in a whisper close to the user's ear, allowing the user to continue the conversation using phrases suggested in real time.

[0757] Step 11:

[0758] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data back to the server.

[0759] Step 12:

[0760] The server feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm.

[0761] Example 1

[0762] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0763] In many modern conversational environments, users face the challenge of generating appropriate phrases in real time and smoothly advancing a conversation. In particular, in business meetings and casual conversations, being able to instantly respond appropriately to the situation requires high communication skills. Therefore, there is a need for technology that allows users to obtain appropriate phrases in a timely manner during a conversation.

[0764] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0765] In this invention, the server includes a means for analyzing speech and converting it into text data, a means for using a generative AI model to generate appropriate phrases based on the text data, and a means for analyzing the user's emotional state, thereby enabling the server to analyze the user's conversation in real time and instantly generate and provide appropriate and humorous phrases.

[0766] "User" refers to an individual who wears an earphone-type device and is the conversation partner.

[0767] "Voice" refers to conversations and other vocal content spoken by a user.

[0768] "Collection means" refers to the equipment and technology used to pick up and appropriately process audio.

[0769] A "terminal" is an earphone-type device worn by a user, and refers to a device that collects, processes, and transmits audio.

[0770] "Server" refers to a computer system for receiving voice data, analyzing it, converting it into text, generating phrases, etc.

[0771] "Generative AI model" refers to an algorithm equipped with artificial intelligence technology that is used to generate appropriate phrases based on text data.

[0772] "Audio playback means" refers to a technique or device that plays back the generated phrase as audio.

[0773] "Natural language processing technology" refers to computer technology for processing, understanding, and generating human language.

[0774] "Emotional state" refers to the user's psychological state, which indicates the tone and emotional state of the conversation.

[0775] The present invention relates to a system in which a user wears an earphone-type device, collects the sound of a conversation in real time, and transmits the collected sound data to a server for analysis. The system of the present invention is implemented as follows.

[0776] 1. The user puts on the device

[0777] When a user puts the earphone-type device in their ear, the device automatically starts up and turns on sound collection mode, allowing it to collect conversations around the user in real time.

[0778] 2. Recording conversations using a device

[0779] The device is equipped with a highly sensitive microphone that collects conversations around the user in real time. The audio data is temporarily stored in the device's local memory, and noise-cancelling technology is used to ensure clear audio. This process uses a microphone with noise-cancelling technology.

[0780] 3. Transmission of audio data by the device

[0781] The device compresses the collected audio data and sends it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end to ensure security. An audio compression codec such as G.711 is used for this compression.

[0782] 4. Receiving and analyzing audio data by the server

[0783] The server receives the voice data sent from the device and uses a speech recognition engine (e.g., Google Speech-to-Text) to decode it. The server converts the voice data into text data, then uses natural language processing (NLP) techniques to analyze the context of the voice and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state.

[0784] 5. Server-generated phrases

[0785] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. This generative AI model has the ability to predict and generate the most appropriate phrases based on the user's speech content and emotional state.

[0786] 6. Sending proposals from the server to the device

[0787] The generated phrase is then compressed again and sent to the device using a low-latency communication protocol. The device then decodes the compressed data and moves on to the next step.

[0788] 7. "Whisper" phrases from your device

[0789] The device converts the received phrase into a voice in real time using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear. This voice playback uses a TTS engine (e.g., Amazon Polly).

[0790] Specific examples

[0791] Example 1: Business meeting

[0792] User: "Today I'd like to talk about a new project."

[0793] The device collects the audio and sends it to the server.

[0794] The server analyzes keywords such as "project" and "want to talk" and performs mood analysis.

[0795] The server uses a "generative AI model" to generate the phrase "Why not talk a little about the progress of the project?" and sends it to the device.

[0796] The terminal whispers, "Why not give us a little update on the progress of the project?"

[0797] Example 2: Casual conversation

[0798] User: "Do you have any plans for the weekend?"

[0799] The device collects the audio and sends it to the server.

[0800] The server analyzes "weekend" and "schedule" and takes into account the user's emotional state.

[0801] The server uses a "generative AI model" to generate the phrase "Why not invite me to a new restaurant that's the talk of the town?" and sends it to the device.

[0802] The device whispers, "Would you like to try a restaurant that's been getting a lot of attention lately?"

[0803] In this way, the present invention provides users with appropriate and humorous phrases in real time, allowing for smooth and engaging conversations.

[0804] Example prompt sentence:

[0805] "When a user says, 'Today I'd like to talk about a new project' in a business meeting, please suggest an appropriate phrase."

[0806] Example output:

[0807] "Maybe we should talk a little bit about how the project is going."

[0808] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0809] Step 1:

[0810] User Wears Device

[0811] The user puts the earphone-type device in their ear. This causes the device to automatically start up and turn on sound collection mode. The input is the state in which the earphone is being worn, and the output is the device switching to sound collection mode.

[0812] Step 2:

[0813] Recording conversations using a device

[0814] The device uses a built-in microphone to collect conversations around the user in real time. The collected audio data is temporarily stored in local memory, and clear audio is obtained using noise-canceling technology. The input is the surrounding audio, and the output is clear audio data that has been subjected to noise-canceling processing. Specifically, the device collects audio every second and stores the data in a buffer.

[0815] Step 3:

[0816] Sending audio data by the device

[0817] The device compresses the collected voice data using a voice compression codec such as G.711 and transmits it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end before transmission. The input is clear voice data, and the output is compressed and encrypted data packets.

[0818] Step 4:

[0819] Receiving and analyzing voice data by the server

[0820] The server receives and decodes the voice data sent from the device. It then uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. It then uses natural language processing (NLP) technology to analyze the context of the text and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state. The input is the decoded voice data, and the output is the analyzed text and emotional information. Specifically, the server receives the voice data, converts it into text through the speech recognition engine, and then analyzes it using an NLP model.

[0821] Step 5:

[0822] Server-generated phrases

[0823] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. The generative AI model predicts and generates the most appropriate phrase based on the prompt sentence. The input is the text data and emotional information, and the output is the generated phrase. Specifically, the server inputs the prompt sentence into the generative AI model and obtains the generated phrase.

[0824] Step 6:

[0825] Sending proposals from the server to the device

[0826] The server compresses the generated phrase and transmits it to the terminal using a low-latency communication protocol. The input is the generated phrase, and the output is a compressed and encrypted data packet. Specifically, the server compresses the generated phrase and encrypts it for transmission as a data packet.

[0827] Step 7:

[0828] "Whisper" phrases from your device

[0829] The device converts the received phrase into audio in real time using Text-to-Speech (TTS) technology and plays it as a "whisper" in the user's ear. The input is encrypted phrase data, and the output is a spoken suggested phrase. Specifically, the device uses a TTS engine to convert the phrase into audio and plays the audio through earphones.

[0830] (Application example 1)

[0831] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0832] In conventional content distribution services, it has been difficult to provide appropriate commentary or supplemental information about the content being viewed in real time. It has also been a challenge to provide accurate information according to the user's conversation and mood. The present invention aims to solve these problems.

[0833] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0834] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for reproducing the phrases in a whisper in the terminal, and means for analyzing the content being viewed and providing commentary and supplemental information in real time, thereby enabling the user to enjoy appropriate commentary and supplemental information in real time even while viewing the content.

[0835] The "means for collecting the user's conversation" is a function for collecting the user's voice using a microphone or the like.

[0836] The "means for transmitting collected conversations to a server" is a function for transmitting collected voice data to a server via a network.

[0837] The "means for analyzing the conversation on the server and generating appropriate phrases" is a function for analyzing the voice data received on the server and generating appropriate phrases based on the context and mood of the conversation.

[0838] The "means for transmitting the generated phrase to the terminal" is a function for transmitting the generated phrase to the terminal via the network.

[0839] The "means for reproducing a phrase in a whisper on the terminal" is a function that uses speech synthesis technology to reproduce the received phrase and whisper it to the user's ear.

[0840] "Means for analyzing content being viewed and providing commentary and supplementary information in real time" refers to a function that analyzes the video or audio content being viewed by the user and generates and provides commentary and supplementary information related to that content in real time.

[0841] The present invention provides a system that collects user conversations and provides appropriate phrases in real time. Specifically, it includes a function that also provides commentary and supplemental information related to the content being viewed in real time. The system of the present invention is implemented as follows.

[0842] 1. Audio collection and transmission

[0843] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user. This audio data is temporarily stored in the device's local memory and undergoes noise cancellation processing. The compressed audio data is then end-to-end encrypted and securely transmitted to a server using a low-latency communication protocol.

[0844] 2. Analysis and phrase generation on the server

[0845] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology. This textualized conversation undergoes contextual analysis and keyword extraction. It also simultaneously analyzes the user's mood, assessing the tone and emotional state of the conversation. Based on this data, appropriate phrases are generated using a generative AI model.

[0846] It also analyzes the content being viewed, analyzing the audio and video of the content and generating relevant commentary and supplemental information. For example, while watching educational content, it can provide academic background information and related new research results in real time.

[0847] 3. Sending and playing phrases

[0848] The generated phrase is then sent back to the device, which then converts the received phrase into voice using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear through the earphone speaker, allowing the user to obtain relevant information in real time while maintaining a natural conversation flow.

[0849] Specific examples

[0850] Situation 1: While watching an educational video

[0851] When a user is watching a scientific video and says, "Next, let's talk about Newtonian mechanics," the device collects this audio and sends it to the server. The server analyzes the content and generates the phrase, "Shall we also talk a bit about the law of gravity?" This allows the user to get additional information about the law of gravity in real time.

[0852] Situation 2: During a casual conversation

[0853] If a user says in a casual conversation, "What are you planning to do this weekend?", the server will generate a phrase like, "Why don't you invite her to try that new restaurant that everyone's talking about?" This allows users to naturally introduce new topics into their conversations.

[0854] Prompt Sentence Examples

[0855] (Example prompt):

[0856] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[0857] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[0858] The present invention enables users to receive appropriate phrases and information in real time that are in line with the content they are viewing or the context of the conversation, thereby realizing a rich communication environment.

[0859] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0860] Step 1:

[0861] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The input is the user's voice, and the output is clear audio data that has been processed with noise cancellation. This audio data is temporarily stored in the device's local memory.

[0862] Step 2:

[0863] The voice data stored on the device is compressed and sent to the server as data packets. The input is clear voice data, and the output is compressed data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted.

[0864] Step 3:

[0865] The server uses natural language processing (NLP) techniques to decode the received voice data and convert it into text. The input is compressed voice data, and the output is the text of the conversation. The server analyzes this text data and performs context analysis and keyword extraction.

[0866] Step 4:

[0867] The server uses a generative AI model to generate appropriate phrases based on the data obtained through context analysis and keyword extraction. The input is the data obtained through context analysis and keyword extraction, and the output is an appropriate phrase. This phrase also takes into account the user's mood.

[0868] Step 5:

[0869] The server sends the generated phrase to the terminal. The input is the generated phrase, and the output is the data sent to the terminal. The data is compressed and sent with the minimum amount of data. A low-latency communication protocol is used for transmission.

[0870] Step 6:

[0871] The device converts the received phrase into voice using Text-to-Speech (TTS) technology. The input is the phrase sent from the server, and the output is voice data. This voice data is played back as a "whisper" in the user's ear through the earphone speaker.

[0872] Step 7:

[0873] The content being viewed is also analyzed, and related commentary and supplementary information are generated. The input is the content data being viewed, and the output is commentary and supplementary information based on the analysis results. This data is generated in real time and provided to the user.

[0874] Example prompt

[0875] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[0876] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[0877] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0878] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then sent back to the terminal and "whispered" to suggest them to the user. The present invention also incorporates an emotion engine that recognizes the user's emotions, and generates appropriate phrases based on the emotional information.

[0879] 1. Recording conversations using a device

[0880] When a user wears the earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory and noise-canceling processing is applied. The audio data is organized into samples at regular time intervals and prepared for transmission to the server.

[0881] 2. Sending conversations by device

[0882] The device compresses the collected audio data and transmits it to the server as data packets. A low-latency communication protocol is used for transmission, maintaining real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[0883] 3. Analysis of conversations by the server

[0884] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques. The textual content undergoes contextual analysis and keyword extraction to identify important keywords and topics. The server also performs mood analysis to assess the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[0885] 4. Emotion Recognition by Emotion Engine

[0886] The server uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's voice characteristics, such as tone, pitch, rhythm, and speed, to determine multiple emotional states. This allows for a precise understanding of the user's current emotions.

[0887] 5. Server-generated phrases

[0888] The server generates appropriate and humorous phrases based on context and emotional information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. Based on a large amount of training data, this algorithm is capable of creating natural-sounding phrases that fit the context. The generated phrases are temporarily stored on the server and immediately sent to the device.

[0889] 6. Sending proposals from the server to the device

[0890] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[0891] 7. "Whisper" phrases from your device

[0892] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[0893] Specific examples

[0894] Example 1: Business meeting

[0895] If a user says, "I'd like to talk about a new project today," the device will pick up this speech and send it to the server. The server will extract the keywords "project" and "want to talk" and use an emotion engine to recognize emotions such as a mixture of tension and excitement. Based on this, the server will generate a phrase such as "It might be helpful if you also mention the progress of the project," and send it to the device. The phrase is suggested to the user in a "whispered" voice.

[0896] Example 2: Casual conversation

[0897] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the emotion of interest. The server generates the phrase "Why don't you try that cafe that's been getting a lot of attention lately?" and suggests it in a whisper from the device.

[0898] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations. By combining it with an emotion engine, more accurate suggestions that are in tune with the user's emotions become possible, improving the quality of communication.

[0899] The processing flow will be explained below.

[0900] Step 1:

[0901] The user wears the earphone-type device. The device uses a built-in microphone to pick up the conversations around the user in real time. The collected audio data is temporarily stored in the device's local memory and noise-canceling processing is applied.

[0902] Step 2:

[0903] The device samples the noise-canceled audio data at regular intervals, compresses it into data packets, and encrypts them. The encrypted audio data is then ready to be sent to the server.

[0904] Step 3:

[0905] The device sends encrypted data packets to the server using a low-latency communication protocol. To maintain real-time performance, the transmitted data is encrypted end-to-end.

[0906] Step 4:

[0907] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology, which is then temporarily stored on the server for analysis.

[0908] Step 5:

[0909] The server performs contextual analysis and keyword extraction based on the text data. Contextual analysis uses grammar rules and statistical models, while keyword extraction uses text mining techniques. After identifying important keywords and topics, the analysis results are passed on to the next processing step.

[0910] Step 6:

[0911] The server uses mood analysis and an emotion engine to recognize the user's emotional state. The emotion engine analyzes characteristics of the voice data, such as pitch, tone, rhythm, and speed, to evaluate the user's emotional state in real time. For example, multiple emotions such as joy, tension, and excitement can be recognized.

[0912] Step 7:

[0913] The server integrates contextual and emotional information to generate appropriate phrases. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. The algorithm is capable of creating natural-sounding phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and are then ready to be sent to the device.

[0914] Step 8:

[0915] The server compresses the generated phrase into a data packet and sends it to the device using a low-latency communication protocol. All data transmission is encrypted and occurs in real time.

[0916] Step 9:

[0917] The device decompresses the received phrase data and converts it into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker, allowing the user to use the suggested phrases in real time.

[0918] Step 10:

[0919] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data to the server.

[0920] Step 11:

[0921] The server then feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm, thus improving the system's ability to consistently deliver optimal phrases.

[0922] Example 2

[0923] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0924] In modern communication, it is important to provide appropriate conversations and responses instantly. However, users often find it difficult to think of appropriate phrases in real time. It is even more difficult to understand the conversation context and the user's emotions and make appropriate suggestions. Therefore, there is a need for a support system that helps users have smooth and engaging conversations.

[0925] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting voice data into text using natural language processing technology, and performing context analysis and keyword extraction, a means for performing emotion analysis, and a means for generating appropriate phrases based on the context and emotion information. This makes it possible to grasp the content of the user's conversation and emotional state in real time and provide appropriate and humorous phrases.

[0926] "User" refers to a person who uses the system to receive conversational assistance.

[0927] "Conversation" refers to verbal communication between a user and another person.

[0928] "Audio equipment" refers to a hardware device for collecting sound, such as an earphone-type device worn by a user.

[0929] "Sound collection" refers to capturing surrounding sounds in real time using audio equipment.

[0930] "Data compression" is the process of reducing the volume of audio data, which is done to increase transmission efficiency.

[0931] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[0932] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language.

[0933] "Text conversion" refers to the process of converting audio data into text data.

[0934] "Contextual analysis" refers to an analytical technique for understanding the content of a conversation and generating appropriate phrases based on that content.

[0935] "Keyword extraction" refers to the process of identifying and extracting important words and phrases from text.

[0936] "Emotion analysis" refers to techniques for identifying a user's emotional state from their vocal characteristics.

[0937] "Phrase generation" refers to the process of creating appropriate sentences based on the results of contextual and sentiment analysis.

[0938] "Data encryption technology" refers to technology that encrypts data to maintain its security.

[0939] A "low latency communication protocol" refers to a communication protocol that minimizes delays in sending and receiving data.

[0940] "Speech synthesis technology" refers to technology for converting text data into voice data.

[0941] "Whisper" refers to a soft voice generated by a terminal, and refers to a voice output format that sounds natural to the user.

[0942] The present invention is a system that uses an earphone-type device worn by the user to collect conversations in real time, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then transmitted back to the terminal and "whispered" to suggest them to the user. This system incorporates an emotion engine that recognizes the user's emotions, and has the advantage of generating appropriate phrases based on the emotion information. The specific steps for implementing the present invention are as follows.

[0943] The user first puts on the earphone-type device. The earphone has a built-in highly sensitive microphone that picks up surrounding conversations in real time. The audio data is temporarily stored in the device's local memory, and noise cancellation processing is performed at the same time. For example, a typical smartphone or portable audio device can be used as the device.

[0944] The device is equipped with data compression software, which compresses the collected audio data. The compressed data is then sent to the server using a low-latency communication protocol, such as WebRTC. The transmitted data is also end-to-end encrypted, protecting the user's privacy.

[0945] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques, possibly using the Google Cloud Speech-to-Text API or similar. The converted text undergoes contextual analysis and keyword extraction to extract important keywords and topics. The server then performs mood analysis to assess the tone and emotional state of the conversation. This analysis can be performed using the DeepAffects API or similar technologies.

[0946] Furthermore, the server is equipped with an emotion engine that recognizes multiple emotional states in real time by analyzing the tone, pitch, rhythm, and speed of the user's voice. Based on this information, it is possible to determine the emotion the user is feeling in the current conversation.

[0947] Based on contextual analysis and sentiment information, the server generates appropriate and humorous phrases using pre-trained generative AI models such as OpenAI's GPT-3, which are then temporarily stored on the server and immediately sent to the device.

[0948] The device converts the received phrase into voice using text-to-speech technology. Mimic3 or a similar voice synthesis engine can be used here. The generated voice data is played back as a "whisper" in the user's ear through the earphone speaker. This "whisper" voice allows the user to instantly use the appropriate phrase.

[0949] Through the above process, the present invention can provide users with appropriate phrases in real time, enabling smooth and engaging conversations in a variety of situations. Furthermore, by combining it with an emotion engine, more accurate suggestions become possible, improving the quality of communication.

[0950] Specific examples

[0951] Example 1: Business meeting

[0952] If a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts keywords like "project" and "want to talk" and uses an emotion engine to recognize tension and excitement. Based on this, the server generates a phrase like, "It might be helpful if you also mention the progress of the project," and sends it to the device. The user receives this suggestion in a "whispered" voice.

[0953] Example 2: Casual conversation

[0954] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the user's interests. The server generates the phrase "Why don't you try that cafe that's been trending lately?" and suggests it in a whisper from the device.

[0955] Prompt Sentence Examples

[0956] "Please suggest how users should talk about projects in business conversations."

[0957] "Show me appropriate phrases to use in casual conversation about weekend plans."

[0958] As a result, the present invention can provide users with appropriate phrases in real time to enable smooth and engaging conversations in a variety of situations.

[0959] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0960] Step 1:

[0961] After the user puts on the earphone-type device, the device activates the built-in microphone to collect surrounding sounds in real time. The input is the user's conversational voice. The device temporarily stores the collected voice data in local memory and performs noise cancellation processing. The output is the voice data after noise cancellation. Specifically, it uses the microphone function of a smartphone or portable audio device.

[0962] Step 2:

[0963] The device processes the noise-canceled audio data by arranging it as samples at regular time intervals and compressing it. The input is the noise-canceled audio data. For example, the G.711 compression format is used for data compression. The compressed data is sent to the server using a low-latency communication protocol (e.g., WebRTC). The output is compressed and encrypted audio data. Again, AES-256 technology is used for end-to-end encryption.

[0964] Step 3:

[0965] The server decodes the received compressed audio data. The input is compressed and encrypted audio data. After decoding, it is converted into PCM data. The server then converts the audio into text using natural language processing (NLP) technology (e.g., Google Cloud Speech-to-Text API). Specifically, the audio signal is analyzed and output as text data. The output is the text of the conversation.

[0966] Step 4:

[0967] The server performs contextual analysis and keyword extraction on the obtained text data. The input is the textual content of the conversation. The NLTK library and other libraries are used for contextual analysis and keyword extraction. As a result of the analysis, important keywords and topics are extracted. The output is keywords and contextual data.

[0968] Step 5:

[0969] The server then performs mood analysis. The input is text data and voice characteristics data. The DeepAffects API is used for emotion analysis. The tone, pitch, rhythm, and speed of the voice are analyzed to determine the user's emotional state. The output is data indicating the user's emotional state (e.g., "tense" or "excited").

[0970] Step 6:

[0971] The server generates appropriate phrases using a generative AI model (e.g., GPT-3) based on contextual analysis and emotional information. The inputs are keywords, contextual data, and emotional data. The generated phrases are temporarily stored on the server. The output is the generated phrase. Here, the behavior of the generative AI model is adjusted by using prompt sentences as examples.

[0972] Step 7:

[0973] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase. The OPUS format is used to compress the data. AES-256 encryption technology is also applied during transmission. The output is the compressed and encrypted phrase data.

[0974] Step 8:

[0975] The device decodes the received phrase and converts it into voice using speech synthesis technology (e.g., Mimic3). The input is compressed and encrypted phrase data. The speech synthesis generates a whispered voice. The generated voice is played to the user's ear through the earphone speaker. The output is a voiced phrase. This "whispered" voice allows the user to use appropriate phrases during actual conversations.

[0976] (Application example 2)

[0977] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0978] Modern autonomous vehicles lack a smooth and comfortable way to communicate with passengers. This can cause stress and anxiety for passengers, potentially reducing the quality of their riding experience. Furthermore, traditional voice assistants lack the ability to adequately analyze emotions and provide prompt and appropriate feedback. This can result in inappropriate responses to passenger instructions and requests.

[0979] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0980] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for playing the phrases in a whisper in the terminal, and means for generating appropriate feedback based on the collected conversation and the analysis results and suggesting the feedback to the user in a voice. This makes it possible to reduce the user's stress and anxiety while riding and provide a comfortable riding experience.

[0981] "User" refers to a person or end user who uses the system.

[0982] "Means for collecting conversation sounds" refers to equipment including a microphone and audio input device for collecting sounds around the user.

[0983] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[0984] "Means for transmitting the conversation to the server" refers to equipment that includes a communication module and protocol for transmitting collected voice data to a remote server over a network.

[0985] "Means for analyzing conversations" refers to natural language processing technologies and algorithms that convert collected audio data into text and analyze its content.

[0986] "Means for generating appropriate phrases" refers to generative AI models and algorithms that generate appropriate responses and suggestions based on analyzed text data and the user's emotional state.

[0987] "Means for transmitting phrases to a terminal" refers to a device that includes a communication module or protocol for transmitting the generated phrases to a user's terminal via a network.

[0988] "Means for playing in a whisper" refers to speakers or earphones that use voice synthesis technology to play phrases received by the terminal to the user at a low volume.

[0989] "Means for generating appropriate feedback and providing audible suggestions to the user" refers to a system in general for providing appropriate feedback and suggestions to the user audibly based on the analysis results.

[0990] To realize the present invention, the following system configuration and processing flow are included.

[0991] The present invention provides a method for collecting and analyzing a user's conversation in real time and suggesting appropriate feedback to the user by utilizing an earphone-type device, a microphone, a server, a communication protocol, and speech synthesis technology.

[0992] Hardware and Software Configuration

[0993] 1. User terminal (earphone-type device)

[0994] Microphone: Picks up the user's conversation and provides noise cancellation.

[0995] Communication module: A module for achieving low latency communication with the server.

[0996] Speaker / Earphone: Phrases received from the server are played back in a whisper using voice synthesis technology.

[0997] 2. Server

[0998] Natural language processing engine (NLP engine): Analyzes voice data and converts it into text.

[0999] Sentiment analysis engine: Analyzes user emotions from voice or text data.

[1000] Generative AI model: Generates appropriate feedback based on the user's conversational content and emotional state.

[1001] Communication module: A module for low-latency, encrypted communication with user devices.

[1002] System processing flow

[1003] 1. Audio collection:

[1004] When a user speaks, the microphone in the earphone-type device picks up the conversation, and the audio data is temporarily stored in the device's local memory and noise-canceling is performed.

[1005] 2. Sending audio data:

[1006] The collected audio data is compressed and encrypted before being sent to a server using a low-latency, highly secure protocol.

[1007] 3. Analysis of audio data:

[1008] The server receives the voice data and converts it into text using a natural language processing engine. It then analyzes the text data to extract important keywords and sentences.

[1009] 4. Emotion analysis:

[1010] In parallel with natural language processing, a sentiment analysis engine assesses the user's emotional state (e.g., happy, sad, nervous) in real time.

[1011] 5. Generate appropriate feedback:

[1012] Based on the analysis results and sentiment data, the generative AI model generates appropriate phrases that include helpful information and suggestions for the user.

[1013] 6. Sending the phrase from the server to the device:

[1014] The generated phrases are compressed and transmitted to the user terminal via a low-latency communication protocol.

[1015] 7. Phrase audio playback:

[1016] The user terminal uses voice synthesis technology to play back the received phrase in a "whispered" voice and suggest it to the user.

[1017] Specific examples

[1018] Example 1: Stress reduction function

[1019] If a passenger says, "I'm feeling a little stressed right now," the voice data is sent to the server. The server's emotion analysis engine detects "stress," and the generative AI model generates a suggestion: "Would you like me to play some music to change your mood?" The words "Would you like me to play some music to change your mood?" are whispered through the earphones.

[1020] Example 2: Tourist information function

[1021] When a passenger asks, "What are the tourist spots around here?", the server analyzes the keywords "tourism" and "spot." The generative AI model generates a suggestion such as, "There's a famous park nearby. Would you like me to show you around?", and a voice whispers through the earphones, "There's a famous park nearby. Would you like me to show you around?"

[1022] Prompt Sentence Examples

[1023] "Generate appropriate suggestions based on passenger statements. Example statement: 'I'm feeling a bit stressed right now.'"

[1024] "Generate suggestions based on the context and emotional state of the utterance. Example utterance: 'What are the tourist attractions around here?'"

[1025] In this way, the present invention provides a system that provides appropriate feedback to the user in real time, thereby realizing a comfortable riding experience for the user.

[1026] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1027] Step 1:

[1028] When a user speaks, the conversation is picked up by a microphone built into the terminal (earphone-type device). The collected audio data undergoes noise cancellation processing and is temporarily stored in local memory. The input is the user's conversational voice, and the output is noise-canceled audio data.

[1029] Step 2:

[1030] The device compresses the noise-canceled voice data and transmits it to the server using a low-latency, encrypted communication protocol. The input is the noise-canceled voice data, and the output is compressed and encrypted voice data packets.

[1031] Step 3:

[1032] The server receives the voice data packets sent from the device, decodes them, and converts them back into the original voice data. It then converts this voice data into text using a natural language processing engine. The input is the encrypted and compressed voice data packets, and the output is the text of the conversation.

[1033] Step 4:

[1034] The server analyzes the textual content of the conversation and extracts context and keywords. It then uses a sentiment analysis engine to evaluate the user's emotional state (e.g., joy, sadness, tension) from the text. The input is the textual content of the conversation, and the output is the extracted keywords and the user's emotional state.

[1035] Step 5:

[1036] The server uses a generative AI model to generate appropriate phrases based on the extracted keywords and emotional state. These generated phrases contain useful information and suggestions for the user. The input is keywords and emotional state, and the output is the generated appropriate phrases.

[1037] Step 6:

[1038] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase and the output is a compressed data packet.

[1039] Step 7:

[1040] The device decodes the data packets received from the server and converts the phrases into audio using text-to-speech technology. It then plays the suggestions to the user in a whispered voice through the earphone speaker. The input is the compressed data packets, and the output is the suggested audio played in a whispered voice.

[1041] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1042] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1043] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1044] [Fourth embodiment]

[1045] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1046] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1047] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1048] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1049] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1050] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1051] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1052] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1053] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1054] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1055] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1056] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1057] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1058] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, transmits the audio data to a server for analysis, and sends appropriate phrases generated by the analysis back to the terminal, where they are "whispered" to suggest the phrases to the user. The present invention is implemented as follows.

[1059] 1. Recording conversations using a device

[1060] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory. Noise cancellation processing is applied to obtain clear audio. The device prepares to send this audio data to a server at regular intervals.

[1061] 2. Sending conversations by device

[1062] The device compresses the voice data and sends it to the server as data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted to ensure security. Data transmission is always in real time, with great care taken to minimize latency.

[1063] 3. Analysis of conversations by the server

[1064] The server decodes the received audio data and converts it into text using natural language processing (NLP). The converted text undergoes contextual analysis and keyword extraction to identify important keywords and topics. Mood analysis is also performed to evaluate the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[1065] 4. Phrase generation by the server

[1066] The server generates appropriate and humorous phrases based on the extracted keywords and mood information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. This algorithm is capable of creating natural phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and immediately sent to the device.

[1067] 5. Proposal transmission from the server to the device

[1068] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[1069] 6. "Whisper" phrases from your device

[1070] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[1071] Specific examples

[1072] Example 1: Business meeting

[1073] When a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts the keywords "project" and "want to talk," performs a mood analysis, and generates an appropriate phrase, such as "Why don't you talk a little about the progress of the project?" The generated phrase is sent to the device and played as a "whisper" in the user's ear.

[1074] Example 2: Casual conversation

[1075] When a user says, "Do you have any plans for the weekend?", the device collects the corresponding voice and sends it to the server. The server analyzes "weekend" and "plans" as keywords and also considers the user's emotional state (e.g., interests), generating the phrase "Why don't you invite me to a new restaurant that's been making waves lately?". This phrase is then whispered through the device: "Do you want to try a restaurant that's been making waves lately?"

[1076] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations.

[1077] The processing flow will be explained below.

[1078] Step 1:

[1079] The user wears the earphone-type device. The device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory.

[1080] Step 2:

[1081] The device applies noise cancellation to the collected voice data to make it clear. The voice data is sampled at regular intervals, compressed into data packets, and encrypted.

[1082] Step 3:

[1083] The device sends compressed and encrypted audio data to the server using a low-latency communication protocol to maintain real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[1084] Step 4:

[1085] The server decodes the received voice data and converts it into text using natural language processing (NLP) techniques.

[1086] Step 5:

[1087] The server performs contextual analysis and keyword extraction on the converted conversation content, identifying important keywords and topics and understanding the context of the conversation.

[1088] Step 6:

[1089] The server performs mood analysis to assess the tone and emotional state of the conversation, thereby capturing the atmosphere of the conversation (e.g., joy, tension, excitement).

[1090] Step 7:

[1091] The server generates appropriate and humorous phrases based on context and mood information, using pre-trained artificial intelligence (AI) algorithms.

[1092] Step 8:

[1093] The server then re-compresses the generated phrases into data packets, which are then sent to the device using a low-latency communication protocol.

[1094] Step 9:

[1095] The device decompresses the received phrase data and converts it from text to voice using text-to-speech technology.

[1096] Step 10:

[1097] The device then plays the generated voice data through the earphone speaker in a whisper close to the user's ear, allowing the user to continue the conversation using phrases suggested in real time.

[1098] Step 11:

[1099] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data back to the server.

[1100] Step 12:

[1101] The server feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm.

[1102] Example 1

[1103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1104] In many modern conversational environments, users face the challenge of generating appropriate phrases in real time and smoothly advancing a conversation. In particular, in business meetings and casual conversations, being able to instantly respond appropriately to the situation requires high communication skills. Therefore, there is a need for technology that allows users to obtain appropriate phrases in a timely manner during a conversation.

[1105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1106] In this invention, the server includes a means for analyzing speech and converting it into text data, a means for using a generative AI model to generate appropriate phrases based on the text data, and a means for analyzing the user's emotional state, thereby enabling the server to analyze the user's conversation in real time and instantly generate and provide appropriate and humorous phrases.

[1107] "User" refers to an individual who wears an earphone-type device and is the conversation partner.

[1108] "Voice" refers to conversations and other vocal content spoken by a user.

[1109] "Collection means" refers to the equipment and technology used to pick up and appropriately process audio.

[1110] A "terminal" is an earphone-type device worn by a user, and refers to a device that collects, processes, and transmits audio.

[1111] "Server" refers to a computer system for receiving voice data, analyzing it, converting it into text, generating phrases, etc.

[1112] "Generative AI model" refers to an algorithm equipped with artificial intelligence technology that is used to generate appropriate phrases based on text data.

[1113] "Audio playback means" refers to a technique or device that plays back the generated phrase as audio.

[1114] "Natural language processing technology" refers to computer technology for processing, understanding, and generating human language.

[1115] "Emotional state" refers to the user's psychological state, which indicates the tone and emotional state of the conversation.

[1116] The present invention relates to a system in which a user wears an earphone-type device, collects the sound of a conversation in real time, and transmits the collected sound data to a server for analysis. The system of the present invention is implemented as follows.

[1117] 1. The user puts on the device

[1118] When a user puts the earphone-type device in their ear, the device automatically starts up and turns on sound collection mode, allowing it to collect conversations around the user in real time.

[1119] 2. Recording conversations using a device

[1120] The device is equipped with a highly sensitive microphone that collects conversations around the user in real time. The audio data is temporarily stored in the device's local memory, and noise-cancelling technology is used to ensure clear audio. This process uses a microphone with noise-cancelling technology.

[1121] 3. Transmission of audio data by the device

[1122] The device compresses the collected audio data and sends it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end to ensure security. An audio compression codec such as G.711 is used for this compression.

[1123] 4. Receiving and analyzing audio data by the server

[1124] The server receives the voice data sent from the device and uses a speech recognition engine (e.g., Google Speech-to-Text) to decode it. The server converts the voice data into text data, then uses natural language processing (NLP) techniques to analyze the context of the voice and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state.

[1125] 5. Server-generated phrases

[1126] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. This generative AI model has the ability to predict and generate the most appropriate phrases based on the user's speech content and emotional state.

[1127] 6. Sending proposals from the server to the device

[1128] The generated phrase is then compressed again and sent to the device using a low-latency communication protocol. The device then decodes the compressed data and moves on to the next step.

[1129] 7. "Whisper" phrases from your device

[1130] The device converts the received phrase into a voice in real time using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear. This voice playback uses a TTS engine (e.g., Amazon Polly).

[1131] Specific examples

[1132] Example 1: Business meeting

[1133] User: "Today I'd like to talk about a new project."

[1134] The device collects the audio and sends it to the server.

[1135] The server analyzes keywords such as "project" and "want to talk" and performs mood analysis.

[1136] The server uses a "generative AI model" to generate the phrase "Why not talk a little about the progress of the project?" and sends it to the device.

[1137] The terminal whispers, "Why not give us a little update on the progress of the project?"

[1138] Example 2: Casual conversation

[1139] User: "Do you have any plans for the weekend?"

[1140] The device collects the audio and sends it to the server.

[1141] The server analyzes "weekend" and "schedule" and takes into account the user's emotional state.

[1142] The server uses a "generative AI model" to generate the phrase "Why not invite me to a new restaurant that's the talk of the town?" and sends it to the device.

[1143] The device whispers, "Would you like to try a restaurant that's been getting a lot of attention lately?"

[1144] In this way, the present invention provides users with appropriate and humorous phrases in real time, allowing for smooth and engaging conversations.

[1145] Example prompt sentence:

[1146] "When a user says, 'Today I'd like to talk about a new project' in a business meeting, please suggest an appropriate phrase."

[1147] Example output:

[1148] "Maybe we should talk a little bit about how the project is going."

[1149] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1150] Step 1:

[1151] User Wears Device

[1152] The user puts the earphone-type device in their ear. This causes the device to automatically start up and turn on sound collection mode. The input is the state in which the earphone is being worn, and the output is the device switching to sound collection mode.

[1153] Step 2:

[1154] Recording conversations using a device

[1155] The device uses a built-in microphone to collect conversations around the user in real time. The collected audio data is temporarily stored in local memory, and clear audio is obtained using noise-canceling technology. The input is the surrounding audio, and the output is clear audio data that has been subjected to noise-canceling processing. Specifically, the device collects audio every second and stores the data in a buffer.

[1156] Step 3:

[1157] Sending audio data by the device

[1158] The device compresses the collected voice data using a voice compression codec such as G.711 and transmits it to the server using a low-latency communication protocol (e.g., WebSocket or UDP). The data is encrypted end-to-end before transmission. The input is clear voice data, and the output is compressed and encrypted data packets.

[1159] Step 4:

[1160] Receiving and analyzing voice data by the server

[1161] The server receives and decodes the voice data sent from the device. It then uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. It then uses natural language processing (NLP) technology to analyze the context of the text and extract keywords. It also uses a sentiment analysis engine to analyze the user's emotional state. The input is the decoded voice data, and the output is the analyzed text and emotional information. Specifically, the server receives the voice data, converts it into text through the speech recognition engine, and then analyzes it using an NLP model.

[1162] Step 5:

[1163] Server-generated phrases

[1164] The server uses a generative AI model (e.g., GPT-3) to generate appropriate and humorous phrases based on the analyzed text data and emotional information. The generative AI model predicts and generates the most appropriate phrase based on the prompt sentence. The input is the text data and emotional information, and the output is the generated phrase. Specifically, the server inputs the prompt sentence into the generative AI model and obtains the generated phrase.

[1165] Step 6:

[1166] Sending proposals from the server to the device

[1167] The server compresses the generated phrase and transmits it to the terminal using a low-latency communication protocol. The input is the generated phrase, and the output is a compressed and encrypted data packet. Specifically, the server compresses the generated phrase and encrypts it for transmission as a data packet.

[1168] Step 7:

[1169] "Whisper" phrases from your device

[1170] The device converts the received phrase into audio in real time using Text-to-Speech (TTS) technology and plays it as a "whisper" in the user's ear. The input is encrypted phrase data, and the output is a spoken suggested phrase. Specifically, the device uses a TTS engine to convert the phrase into audio and plays the audio through earphones.

[1171] (Application example 1)

[1172] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1173] In conventional content distribution services, it has been difficult to provide appropriate commentary or supplemental information about the content being viewed in real time. It has also been a challenge to provide accurate information according to the user's conversation and mood. The present invention aims to solve these problems.

[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1175] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for reproducing the phrases in a whisper in the terminal, and means for analyzing the content being viewed and providing commentary and supplemental information in real time, thereby enabling the user to enjoy appropriate commentary and supplemental information in real time even while viewing the content.

[1176] The "means for collecting the user's conversation" is a function for collecting the user's voice using a microphone or the like.

[1177] The "means for transmitting collected conversations to a server" is a function for transmitting collected voice data to a server via a network.

[1178] The "means for analyzing the conversation on the server and generating appropriate phrases" is a function for analyzing the voice data received on the server and generating appropriate phrases based on the context and mood of the conversation.

[1179] The "means for transmitting the generated phrase to the terminal" is a function for transmitting the generated phrase to the terminal via the network.

[1180] The "means for reproducing a phrase in a whisper on the terminal" is a function that uses speech synthesis technology to reproduce the received phrase and whisper it to the user's ear.

[1181] "Means for analyzing content being viewed and providing commentary and supplementary information in real time" refers to a function that analyzes the video or audio content being viewed by the user and generates and provides commentary and supplementary information related to that content in real time.

[1182] The present invention provides a system that collects user conversations and provides appropriate phrases in real time. Specifically, it includes a function that also provides commentary and supplemental information related to the content being viewed in real time. The system of the present invention is implemented as follows.

[1183] 1. Audio collection and transmission

[1184] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user. This audio data is temporarily stored in the device's local memory and undergoes noise cancellation processing. The compressed audio data is then end-to-end encrypted and securely transmitted to a server using a low-latency communication protocol.

[1185] 2. Analysis and phrase generation on the server

[1186] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology. This textualized conversation undergoes contextual analysis and keyword extraction. It also simultaneously analyzes the user's mood, assessing the tone and emotional state of the conversation. Based on this data, appropriate phrases are generated using a generative AI model.

[1187] It also analyzes the content being viewed, analyzing the audio and video of the content and generating relevant commentary and supplemental information. For example, while watching educational content, it can provide academic background information and related new research results in real time.

[1188] 3. Sending and playing phrases

[1189] The generated phrase is then sent back to the device, which then converts the received phrase into voice using Text-to-Speech (TTS) technology and plays it back as a whisper in the user's ear through the earphone speaker, allowing the user to obtain relevant information in real time while maintaining a natural conversation flow.

[1190] Specific examples

[1191] Situation 1: While watching an educational video

[1192] When a user is watching a scientific video and says, "Next, let's talk about Newtonian mechanics," the device collects this audio and sends it to the server. The server analyzes the content and generates the phrase, "Shall we also talk a bit about the law of gravity?" This allows the user to get additional information about the law of gravity in real time.

[1193] Situation 2: During a casual conversation

[1194] If a user says in a casual conversation, "What are you planning to do this weekend?", the server will generate a phrase like, "Why don't you invite her to try that new restaurant that everyone's talking about?" This allows users to naturally introduce new topics into their conversations.

[1195] Prompt Sentence Examples

[1196] (Example prompt):

[1197] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[1198] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[1199] The present invention enables users to receive appropriate phrases and information in real time that are in line with the content they are viewing or the context of the conversation, thereby realizing a rich communication environment.

[1200] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1201] Step 1:

[1202] When a user wears an earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The input is the user's voice, and the output is clear audio data that has been processed with noise cancellation. This audio data is temporarily stored in the device's local memory.

[1203] Step 2:

[1204] The voice data stored on the device is compressed and sent to the server as data packets. The input is clear voice data, and the output is compressed data packets. A low-latency communication protocol is used for transmission, and the voice data is end-to-end encrypted.

[1205] Step 3:

[1206] The server uses natural language processing (NLP) techniques to decode the received voice data and convert it into text. The input is compressed voice data, and the output is the text of the conversation. The server analyzes this text data and performs context analysis and keyword extraction.

[1207] Step 4:

[1208] The server uses a generative AI model to generate appropriate phrases based on the data obtained through context analysis and keyword extraction. The input is the data obtained through context analysis and keyword extraction, and the output is an appropriate phrase. This phrase also takes into account the user's mood.

[1209] Step 5:

[1210] The server sends the generated phrase to the terminal. The input is the generated phrase, and the output is the data sent to the terminal. The data is compressed and sent with the minimum amount of data. A low-latency communication protocol is used for transmission.

[1211] Step 6:

[1212] The device converts the received phrase into voice using Text-to-Speech (TTS) technology. The input is the phrase sent from the server, and the output is voice data. This voice data is played back as a "whisper" in the user's ear through the earphone speaker.

[1213] Step 7:

[1214] The content being viewed is also analyzed, and related commentary and supplementary information are generated. The input is the content data being viewed, and the output is commentary and supplementary information based on the analysis results. This data is generated in real time and provided to the user.

[1215] Example prompt

[1216] Analyze user speech: The user says, "Next, let's talk about Newtonian mechanics." Generate the appropriate phrase.

[1217] Phrase Generation: Shall we also touch on the laws of gravity for a moment?

[1218] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1219] The present invention provides a system that collects conversations in real time using an earphone-type device worn by the user, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then sent back to the terminal and "whispered" to suggest them to the user. The present invention also incorporates an emotion engine that recognizes the user's emotions, and generates appropriate phrases based on the emotional information.

[1220] 1. Recording conversations using a device

[1221] When a user wears the earphone-type device, the device uses a built-in microphone to pick up conversations around the user in real time. The audio data is temporarily stored in the device's local memory and noise-canceling processing is applied. The audio data is organized into samples at regular time intervals and prepared for transmission to the server.

[1222] 2. Sending conversations by device

[1223] The device compresses the collected audio data and transmits it to the server as data packets. A low-latency communication protocol is used for transmission, maintaining real-time performance. The transmitted data is end-to-end encrypted to protect privacy.

[1224] 3. Analysis of conversations by the server

[1225] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques. The textual content undergoes contextual analysis and keyword extraction to identify important keywords and topics. The server also performs mood analysis to assess the tone and emotional state of the conversation (e.g., joy, tension, excitement).

[1226] 4. Emotion Recognition by Emotion Engine

[1227] The server uses an emotion engine to recognize the user's emotional state in real time. The emotion engine analyzes the user's voice characteristics, such as tone, pitch, rhythm, and speed, to determine multiple emotional states. This allows for a precise understanding of the user's current emotions.

[1228] 5. Server-generated phrases

[1229] The server generates appropriate and humorous phrases based on context and emotional information. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. Based on a large amount of training data, this algorithm is capable of creating natural-sounding phrases that fit the context. The generated phrases are temporarily stored on the server and immediately sent to the device.

[1230] 6. Sending proposals from the server to the device

[1231] The server compresses the generated phrases to minimize data volume when sending them to the terminal, and uses a low-latency communication protocol to provide immediate feedback during user discussions.

[1232] 7. "Whisper" phrases from your device

[1233] The device converts the received phrase into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker. This process allows the user to use the appropriate phrase on the spot.

[1234] Specific examples

[1235] Example 1: Business meeting

[1236] If a user says, "I'd like to talk about a new project today," the device will pick up this speech and send it to the server. The server will extract the keywords "project" and "want to talk" and use an emotion engine to recognize emotions such as a mixture of tension and excitement. Based on this, the server will generate a phrase such as "It might be helpful if you also mention the progress of the project," and send it to the device. The phrase is suggested to the user in a "whispered" voice.

[1237] Example 2: Casual conversation

[1238] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the emotion of interest. The server generates the phrase "Why don't you try that cafe that's been getting a lot of attention lately?" and suggests it in a whisper from the device.

[1239] As a result, the present invention provides users with appropriate and humorous phrases in real time, enabling smooth and engaging conversations in a variety of situations. By combining it with an emotion engine, more accurate suggestions that are in tune with the user's emotions become possible, improving the quality of communication.

[1240] The processing flow will be explained below.

[1241] Step 1:

[1242] The user wears the earphone-type device. The device uses a built-in microphone to pick up the conversations around the user in real time. The collected audio data is temporarily stored in the device's local memory and noise-canceling processing is applied.

[1243] Step 2:

[1244] The device samples the noise-canceled audio data at regular intervals, compresses it into data packets, and encrypts them. The encrypted audio data is then ready to be sent to the server.

[1245] Step 3:

[1246] The device sends encrypted data packets to the server using a low-latency communication protocol. To maintain real-time performance, the transmitted data is encrypted end-to-end.

[1247] Step 4:

[1248] The server decodes the received voice data and converts it into text using natural language processing (NLP) technology, which is then temporarily stored on the server for analysis.

[1249] Step 5:

[1250] The server performs contextual analysis and keyword extraction based on the text data. Contextual analysis uses grammar rules and statistical models, while keyword extraction uses text mining techniques. After identifying important keywords and topics, the analysis results are passed on to the next processing step.

[1251] Step 6:

[1252] The server uses mood analysis and an emotion engine to recognize the user's emotional state. The emotion engine analyzes characteristics of the voice data, such as pitch, tone, rhythm, and speed, to evaluate the user's emotional state in real time. For example, multiple emotions such as joy, tension, and excitement can be recognized.

[1253] Step 7:

[1254] The server integrates contextual and emotional information to generate appropriate phrases. This generation is performed using a pre-trained artificial intelligence (AI) algorithm. The algorithm is capable of creating natural-sounding phrases that fit the context based on a large amount of training data. The generated phrases are temporarily stored on the server and are then ready to be sent to the device.

[1255] Step 8:

[1256] The server compresses the generated phrase into a data packet and sends it to the device using a low-latency communication protocol. All data transmission is encrypted and occurs in real time.

[1257] Step 9:

[1258] The device decompresses the received phrase data and converts it into voice using text-to-speech technology. The generated voice data is then played back as a whisper in the user's ear through the earphone speaker, allowing the user to use the suggested phrases in real time.

[1259] Step 10:

[1260] The user can choose whether to actually use the suggested phrases. The device monitors the user's reactions and usage of the phrases, and reports this data to the server.

[1261] Step 11:

[1262] The server then feeds the reported data into a machine learning feedback loop to help improve the phrase generation algorithm, thus improving the system's ability to consistently deliver optimal phrases.

[1263] Example 2

[1264] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1265] In modern communication, it is important to provide appropriate conversations and responses instantly. However, users often find it difficult to think of appropriate phrases in real time. It is even more difficult to understand the conversation context and the user's emotions and make appropriate suggestions. Therefore, there is a need for a support system that helps users have smooth and engaging conversations.

[1266] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting voice data into text using natural language processing technology, and performing context analysis and keyword extraction, a means for performing emotion analysis, and a means for generating appropriate phrases based on the context and emotion information. This makes it possible to grasp the content of the user's conversation and emotional state in real time and provide appropriate and humorous phrases.

[1267] "User" refers to a person who uses the system to receive conversational assistance.

[1268] "Conversation" refers to verbal communication between a user and another person.

[1269] "Audio equipment" refers to a hardware device for collecting sound, such as an earphone-type device worn by a user.

[1270] "Sound collection" refers to capturing surrounding sounds in real time using audio equipment.

[1271] "Data compression" is the process of reducing the volume of audio data, which is done to increase transmission efficiency.

[1272] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[1273] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language.

[1274] "Text conversion" refers to the process of converting audio data into text data.

[1275] "Contextual analysis" refers to an analytical technique for understanding the content of a conversation and generating appropriate phrases based on that content.

[1276] "Keyword extraction" refers to the process of identifying and extracting important words and phrases from text.

[1277] "Emotion analysis" refers to techniques for identifying a user's emotional state from their vocal characteristics.

[1278] "Phrase generation" refers to the process of creating appropriate sentences based on the results of contextual and sentiment analysis.

[1279] "Data encryption technology" refers to technology that encrypts data to maintain its security.

[1280] A "low latency communication protocol" refers to a communication protocol that minimizes delays in sending and receiving data.

[1281] "Speech synthesis technology" refers to technology for converting text data into voice data.

[1282] "Whisper" refers to a soft voice generated by a terminal, and refers to a voice output format that sounds natural to the user.

[1283] The present invention is a system that uses an earphone-type device worn by the user to collect conversations in real time, and transmits the audio data to a server for analysis. Appropriate phrases generated through the analysis are then transmitted back to the terminal and "whispered" to suggest them to the user. This system incorporates an emotion engine that recognizes the user's emotions, and has the advantage of generating appropriate phrases based on the emotion information. The specific steps for implementing the present invention are as follows.

[1284] The user first puts on the earphone-type device. The earphone has a built-in highly sensitive microphone that picks up surrounding conversations in real time. The audio data is temporarily stored in the device's local memory, and noise cancellation processing is performed at the same time. For example, a typical smartphone or portable audio device can be used as the device.

[1285] The device is equipped with data compression software, which compresses the collected audio data. The compressed data is then sent to the server using a low-latency communication protocol, such as WebRTC. The transmitted data is also end-to-end encrypted, protecting the user's privacy.

[1286] The server decodes the received audio data and converts it to text using natural language processing (NLP) techniques, possibly using the Google Cloud Speech-to-Text API or similar. The converted text undergoes contextual analysis and keyword extraction to extract important keywords and topics. The server then performs mood analysis to assess the tone and emotional state of the conversation. This analysis can be performed using the DeepAffects API or similar technologies.

[1287] Furthermore, the server is equipped with an emotion engine that recognizes multiple emotional states in real time by analyzing the tone, pitch, rhythm, and speed of the user's voice. Based on this information, it is possible to determine the emotion the user is feeling in the current conversation.

[1288] Based on contextual analysis and sentiment information, the server generates appropriate and humorous phrases using pre-trained generative AI models such as OpenAI's GPT-3, which are then temporarily stored on the server and immediately sent to the device.

[1289] The device converts the received phrase into voice using text-to-speech technology. Mimic3 or a similar voice synthesis engine can be used here. The generated voice data is played back as a "whisper" in the user's ear through the earphone speaker. This "whisper" voice allows the user to instantly use the appropriate phrase.

[1290] Through the above process, the present invention can provide users with appropriate phrases in real time, enabling smooth and engaging conversations in a variety of situations. Furthermore, by combining it with an emotion engine, more accurate suggestions become possible, improving the quality of communication.

[1291] Specific examples

[1292] Example 1: Business meeting

[1293] If a user says, "I'd like to talk about a new project today," the device picks up this speech and sends it to the server. The server extracts keywords like "project" and "want to talk" and uses an emotion engine to recognize tension and excitement. Based on this, the server generates a phrase like, "It might be helpful if you also mention the progress of the project," and sends it to the device. The user receives this suggestion in a "whispered" voice.

[1294] Example 2: Casual conversation

[1295] When a user says, "Do you have any plans for the weekend?", the device collects the voice and sends it to the server. The server analyzes the keywords "weekend" and "plans," and the emotion engine recognizes the user's interests. The server generates the phrase "Why don't you try that cafe that's been trending lately?" and suggests it in a whisper from the device.

[1296] Prompt Sentence Examples

[1297] "Please suggest how users should talk about projects in business conversations."

[1298] "Show me appropriate phrases to use in casual conversation about weekend plans."

[1299] As a result, the present invention can provide users with appropriate phrases in real time to enable smooth and engaging conversations in a variety of situations.

[1300] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1301] Step 1:

[1302] After the user puts on the earphone-type device, the device activates the built-in microphone to collect surrounding sounds in real time. The input is the user's conversational voice. The device temporarily stores the collected voice data in local memory and performs noise cancellation processing. The output is the voice data after noise cancellation. Specifically, it uses the microphone function of a smartphone or portable audio device.

[1303] Step 2:

[1304] The device processes the noise-canceled audio data by arranging it as samples at regular time intervals and compressing it. The input is the noise-canceled audio data. For example, the G.711 compression format is used for data compression. The compressed data is sent to the server using a low-latency communication protocol (e.g., WebRTC). The output is compressed and encrypted audio data. Again, AES-256 technology is used for end-to-end encryption.

[1305] Step 3:

[1306] The server decodes the received compressed audio data. The input is compressed and encrypted audio data. After decoding, it is converted into PCM data. The server then converts the audio into text using natural language processing (NLP) technology (e.g., Google Cloud Speech-to-Text API). Specifically, the audio signal is analyzed and output as text data. The output is the text of the conversation.

[1307] Step 4:

[1308] The server performs contextual analysis and keyword extraction on the obtained text data. The input is the textual content of the conversation. The NLTK library and other libraries are used for contextual analysis and keyword extraction. As a result of the analysis, important keywords and topics are extracted. The output is keywords and contextual data.

[1309] Step 5:

[1310] The server then performs mood analysis. The input is text data and voice characteristics data. The DeepAffects API is used for emotion analysis. The tone, pitch, rhythm, and speed of the voice are analyzed to determine the user's emotional state. The output is data indicating the user's emotional state (e.g., "tense" or "excited").

[1311] Step 6:

[1312] The server generates appropriate phrases using a generative AI model (e.g., GPT-3) based on contextual analysis and emotional information. The inputs are keywords, contextual data, and emotional data. The generated phrases are temporarily stored on the server. The output is the generated phrase. Here, the behavior of the generative AI model is adjusted by using prompt sentences as examples.

[1313] Step 7:

[1314] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase. The OPUS format is used to compress the data. AES-256 encryption technology is also applied during transmission. The output is the compressed and encrypted phrase data.

[1315] Step 8:

[1316] The device decodes the received phrase and converts it into voice using speech synthesis technology (e.g., Mimic3). The input is compressed and encrypted phrase data. The speech synthesis generates a whispered voice. The generated voice is played to the user's ear through the earphone speaker. The output is a voiced phrase. This "whispered" voice allows the user to use appropriate phrases during actual conversations.

[1317] (Application example 2)

[1318] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1319] Modern autonomous vehicles lack a smooth and comfortable way to communicate with passengers. This can cause stress and anxiety for passengers, potentially reducing the quality of their riding experience. Furthermore, traditional voice assistants lack the ability to adequately analyze emotions and provide prompt and appropriate feedback. This can result in inappropriate responses to passenger instructions and requests.

[1320] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1321] In this invention, the server includes means for collecting the user's conversation, means for transmitting the collected conversation to the server, means for analyzing the conversation in the server and generating appropriate phrases, means for transmitting the generated phrases to the terminal, means for playing the phrases in a whisper in the terminal, and means for generating appropriate feedback based on the collected conversation and the analysis results and suggesting the feedback to the user in a voice. This makes it possible to reduce the user's stress and anxiety while riding and provide a comfortable riding experience.

[1322] "User" refers to a person or end user who uses the system.

[1323] "Means for collecting conversation sounds" refers to equipment including a microphone and audio input device for collecting sounds around the user.

[1324] "Server" refers to a remote computer system that analyzes speech data and generates phrases.

[1325] "Means for transmitting the conversation to the server" refers to equipment that includes a communication module and protocol for transmitting collected voice data to a remote server over a network.

[1326] "Means for analyzing conversations" refers to natural language processing technologies and algorithms that convert collected audio data into text and analyze its content.

[1327] "Means for generating appropriate phrases" refers to generative AI models and algorithms that generate appropriate responses and suggestions based on analyzed text data and the user's emotional state.

[1328] "Means for transmitting phrases to a terminal" refers to a device that includes a communication module or protocol for transmitting the generated phrases to a user's terminal via a network.

[1329] "Means for playing in a whisper" refers to speakers or earphones that use voice synthesis technology to play phrases received by the terminal to the user at a low volume.

[1330] "Means for generating appropriate feedback and providing audible suggestions to the user" refers to a system in general for providing appropriate feedback and suggestions to the user audibly based on the analysis results.

[1331] To realize the present invention, the following system configuration and processing flow are included.

[1332] The present invention provides a method for collecting and analyzing a user's conversation in real time and suggesting appropriate feedback to the user by utilizing an earphone-type device, a microphone, a server, a communication protocol, and speech synthesis technology.

[1333] Hardware and Software Configuration

[1334] 1. User terminal (earphone-type device)

[1335] Microphone: Picks up the user's conversation and provides noise cancellation.

[1336] Communication module: A module for achieving low latency communication with the server.

[1337] Speaker / Earphone: Phrases received from the server are played back in a whisper using voice synthesis technology.

[1338] 2. Server

[1339] Natural language processing engine (NLP engine): Analyzes voice data and converts it into text.

[1340] Sentiment analysis engine: Analyzes user emotions from voice or text data.

[1341] Generative AI model: Generates appropriate feedback based on the user's conversational content and emotional state.

[1342] Communication module: A module for low-latency, encrypted communication with user devices.

[1343] System processing flow

[1344] 1. Audio collection:

[1345] When a user speaks, the microphone in the earphone-type device picks up the conversation, and the audio data is temporarily stored in the device's local memory and noise-canceling is performed.

[1346] 2. Sending audio data:

[1347] The collected audio data is compressed and encrypted before being sent to a server using a low-latency, highly secure protocol.

[1348] 3. Analysis of audio data:

[1349] The server receives the voice data and converts it into text using a natural language processing engine. It then analyzes the text data to extract important keywords and sentences.

[1350] 4. Emotion analysis:

[1351] In parallel with natural language processing, a sentiment analysis engine assesses the user's emotional state (e.g., happy, sad, nervous) in real time.

[1352] 5. Generate appropriate feedback:

[1353] Based on the analysis results and sentiment data, the generative AI model generates appropriate phrases that include helpful information and suggestions for the user.

[1354] 6. Sending the phrase from the server to the device:

[1355] The generated phrases are compressed and transmitted to the user terminal via a low-latency communication protocol.

[1356] 7. Phrase audio playback:

[1357] The user terminal uses voice synthesis technology to play back the received phrase in a "whispered" voice and suggest it to the user.

[1358] Specific examples

[1359] Example 1: Stress reduction function

[1360] If a passenger says, "I'm feeling a little stressed right now," the voice data is sent to the server. The server's emotion analysis engine detects "stress," and the generative AI model generates a suggestion: "Would you like me to play some music to change your mood?" The words "Would you like me to play some music to change your mood?" are whispered through the earphones.

[1361] Example 2: Tourist information function

[1362] When a passenger asks, "What are the tourist spots around here?", the server analyzes the keywords "tourism" and "spot." The generative AI model generates a suggestion such as, "There's a famous park nearby. Would you like me to show you around?", and a voice whispers through the earphones, "There's a famous park nearby. Would you like me to show you around?"

[1363] Prompt Sentence Examples

[1364] "Generate appropriate suggestions based on passenger statements. Example statement: 'I'm feeling a bit stressed right now.'"

[1365] "Generate suggestions based on the context and emotional state of the utterance. Example utterance: 'What are the tourist attractions around here?'"

[1366] In this way, the present invention provides a system that provides appropriate feedback to the user in real time, thereby realizing a comfortable riding experience for the user.

[1367] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1368] Step 1:

[1369] When a user speaks, the conversation is picked up by a microphone built into the terminal (earphone-type device). The collected audio data undergoes noise cancellation processing and is temporarily stored in local memory. The input is the user's conversational voice, and the output is noise-canceled audio data.

[1370] Step 2:

[1371] The device compresses the noise-canceled voice data and transmits it to the server using a low-latency, encrypted communication protocol. The input is the noise-canceled voice data, and the output is compressed and encrypted voice data packets.

[1372] Step 3:

[1373] The server receives the voice data packets sent from the device, decodes them, and converts them back into the original voice data. It then converts this voice data into text using a natural language processing engine. The input is the encrypted and compressed voice data packets, and the output is the text of the conversation.

[1374] Step 4:

[1375] The server analyzes the textual content of the conversation and extracts context and keywords. It then uses a sentiment analysis engine to evaluate the user's emotional state (e.g., joy, sadness, tension) from the text. The input is the textual content of the conversation, and the output is the extracted keywords and the user's emotional state.

[1376] Step 5:

[1377] The server uses a generative AI model to generate appropriate phrases based on the extracted keywords and emotional state. These generated phrases contain useful information and suggestions for the user. The input is keywords and emotional state, and the output is the generated appropriate phrases.

[1378] Step 6:

[1379] The server compresses the generated phrase and sends it to the terminal using a low-latency communication protocol. The input is the generated phrase and the output is a compressed data packet.

[1380] Step 7:

[1381] The device decodes the data packets received from the server and converts the phrases into audio using text-to-speech technology. It then plays the suggestions to the user in a whispered voice through the earphone speaker. The input is the compressed data packets, and the output is the suggested audio played in a whispered voice.

[1382] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1383] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1384] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1385] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1386] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1387] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1388] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1389] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1390] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1391] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1392] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1393] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1394] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1395] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1396] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1397] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1398] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1399] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1400] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1401] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1402] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1403] The following is further disclosed regarding the above embodiment.

[1404] (Claim 1)

[1405] A means for collecting user conversations;

[1406] means for transmitting the collected conversation to a server;

[1407] A means for analyzing the conversation in the server and generating appropriate phrases;

[1408] means for transmitting the generated phrase to a terminal;

[1409] means for whispering the phrase at the terminal;

[1410] A system including:

[1411] (Claim 2)

[1412] 10. The system of claim 1, wherein the conversation analysis uses natural language processing techniques.

[1413] (Claim 3)

[1414] 10. The system of claim 1, further comprising means for analyzing a user's mood.

[1415] (Claim 4)

[1416] The system according to claim 3, wherein phrases are generated based on the results of the mood analysis.

[1417] (Claim 5)

[1418] 10. The system of claim 1, which uses an artificial intelligence algorithm for phrase generation.

[1419] (Claim 6)

[1420] 10. The system of claim 1, further comprising means for real-time, low-latency communication.

[1421] (Claim 7)

[1422] 10. The system of claim 1, wherein the phrase is reproduced in a whisper using speech synthesis technology.

[1423] "Example 1"

[1424] (Claim 1)

[1425] means for collecting user voice;

[1426] a terminal for processing the collected audio;

[1427] means for transmitting the processed audio to a server;

[1428] A means for analyzing the voice and converting it into text data in a server;

[1429] using a generative AI model to generate appropriate phrases based on the text data; and

[1430] means for transmitting the generated phrase to a terminal;

[1431] means for conveying the phrase generated in the terminal to the user by means of voice reproduction means;

[1432] A system including:

[1433] (Claim 2)

[1434] 10. The system of claim 1, wherein the speech analysis uses natural language processing techniques.

[1435] (Claim 3)

[1436] 10. The system of claim 1, further comprising means for analyzing the emotional state of the user at the server.

[1437] "Application Example 1"

[1438] (Claim 1)

[1439] A means for collecting user conversations;

[1440] means for transmitting the collected conversation to a server;

[1441] A means for analyzing the conversation in the server and generating appropriate phrases;

[1442] means for transmitting the generated phrase to a terminal;

[1443] means for whispering the phrase at the terminal;

[1444] A means to analyze the content being viewed and provide commentary and supplementary information in real time,

[1445] A system including:

[1446] (Claim 2)

[1447] 10. The system of claim 1, wherein natural language processing techniques are used in conversation analysis and content analysis.

[1448] (Claim 3)

[1449] 10. The system of claim 1, further comprising means for analyzing a user's mood and generating optimal phrases related to content viewing.

[1450] "Example 2: Combining Emotion Engines"

[1451] (Claim 1)

[1452] A means for collecting user conversations with an audio device;

[1453] A means for compressing the collected conversation data and transmitting it to a server;

[1454] A server converts the voice data into text using natural language processing technology, and performs context analysis and keyword extraction;

[1455] means for performing emotion analysis in the server;

[1456] a means for generating appropriate phrases based on context and sentiment information;

[1457] means for transmitting the generated phrase to a terminal;

[1458] a means for playing the phrase as a whisper on the device using speech synthesis technology;

[1459] A system including:

[1460] (Claim 2)

[1461] 10. The system of claim 1, further comprising means at the server for determining the emotional state of the user using emotion recognition technology.

[1462] (Claim 3)

[1463] 10. The system of claim 1, further comprising means for transmitting and receiving data using a low latency communication protocol and data encryption techniques.

[1464] "Application example 2 when combining emotion engines"

[1465] (Claim 1)

[1466] A means for collecting user conversations;

[1467] means for transmitting the collected conversation to a server;

[1468] A means for analyzing the conversation in the server and generating appropriate phrases;

[1469] means for transmitting the generated phrase to a terminal;

[1470] means for whispering the phrase at the terminal;

[1471] A means for generating appropriate feedback based on collected conversations and analysis results and providing the feedback to the user by voice;

[1472] A system including:

[1473] (Claim 2)

[1474] 10. The system of claim 1, wherein the conversation analysis uses natural language processing techniques.

[1475] (Claim 3)

[1476] 10. The system of claim 1, further comprising means for analyzing a user's mood. [Explanation of symbols]

[1477] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for collecting user conversations; means for transmitting the collected conversation to a server; A means for analyzing the conversation in the server and generating appropriate phrases; means for transmitting the generated phrase to a terminal; means for whispering the phrase at the terminal; A system including:

2. The system of claim 1 , wherein the conversation analysis uses natural language processing techniques.

3. 10. The system of claim 1, further comprising means for analyzing a user's mood.

4. The system according to claim 3, wherein phrases are generated based on the results of the mood analysis.

5. 10. The system of claim 1, wherein the system uses an artificial intelligence algorithm for phrase generation.

6. 10. The system of claim 1, further comprising means for real-time, low-latency communication.

7. 10. The system of claim 1, wherein the phrases are reproduced in a whisper using speech synthesis technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A