system

JP2026085729APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

In online meetings and lectures, participants face psychological barriers to expressing their opinions, leading to reduced communication quality and anxiety for lecturers, as they struggle to gauge understanding and interest levels.

Method used

A system that acquires speech information, converts it to text, analyzes the context, and generates appropriate responses displayed in real-time to facilitate participant engagement and feedback.

Benefits of technology

Enhances communication quality by allowing participants to naturally contribute and receive feedback, enabling lecturers to adjust their content in real-time, thus improving the learning environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085729000001_ABST
    Figure 2026085729000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for acquiring speech information, A conversion means for converting the acquired speech information into speech text, A generation means that analyzes the context based on the converted speech text and generates an appropriate response, A display means for displaying the generated response on the user interface, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In online meetings and lectures, there is a psychological hurdle for listeners to state their opinions, and as a result, it is difficult for the lecturer to grasp the understanding and interest levels of the participants. Such a situation reduces the quality of meetings and lectures and causes anxiety for the lecturer, so a mechanism for smoothing communication between both participants and lecturers is required

Means for Solving the Problems

[0005] This system facilitates real-time communication in online meetings and lectures by acquiring speech information, converting it to speech-to-text, analyzing the context based on the converted text, and generating appropriate responses. Furthermore, by providing a mechanism to display the generated responses in the user interface, participants' contributions are naturally elicited, and feedback can be obtained on whether the speaker's presentation content is being properly conveyed to the participants. As a result, the quality of the learning environment can be improved.

[0006] "Speech information" refers to audio data and associated information spoken by participants or speakers during online meetings and lectures.

[0007] "Acquisition means" refers to the function of equipment or software that collects speech information from an online platform and converts it into a format that can be processed within the system.

[0008] "Speech-to-text" refers to text data obtained by converting speech data into text information using speech recognition technology.

[0009] "Conversion means" refers to a device or program that performs a process to convert input audio information into text format.

[0010] "Context" refers to a series of background information and sequences, including the scene or situation in which a particular statement or content is used.

[0011] "Generation means" refers to a function that uses machine learning models and algorithms to automatically create appropriate responses based on the analyzed context.

[0012] "Response" refers to text or audio information that includes appropriate reactions and comments to participants or speakers, generated based on the analyzed context.

[0013] "Display means" refers to a part of a system that displays the generated response on a user interface, making it easily accessible to participants.

[0014] "User interface" refers to the means of interaction, such as screen displays and audio outputs, that allow the user of a system to receive information visually or aurally. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor. <0000​​​

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] In one embodiment of the invention, a system is constructed to improve real-time communication in online meetings and lectures. The specific operation of this system is described below.

[0037] First, the server acquires speech information in real time from the online meeting platform. This information includes voice data and speaker IDs, and is continuously collected as the meeting progresses.

[0038] Next, the server converts the acquired audio data into text format using speech recognition technology. This converted audio-text becomes the basis for analyzing the content of the conversation. For example, if the speaker is discussing a new product, relevant keywords and sentences will be reflected in the text.

[0039] Next, the server's generation system analyzes the context using the generated text and produces appropriate responses and comments. This generation process utilizes machine learning models to generate natural interjections and questions that fit the flow of the conversation. For example, if the speaker is talking about a new feature, the AI ​​might create a response such as, "How will that feature improve our work?"

[0040] The generated response is immediately sent from the server to the terminal. On the terminal, this response is displayed in a chat format via a user interface, allowing participants to review it and enter their own opinions or questions as needed. This user interface is designed with ease of use and visibility in mind, making it intuitive for participants to use.

[0041] Furthermore, by allowing users to participate in the discussion using interjections and comments provided through this system, speakers can instantly grasp the audience's interest and level of understanding. This enables speakers to adjust the information they provide in real time, which is expected to improve the quality of online meetings.

[0042] By utilizing such a system, smooth communication can be promoted for both speakers and listeners, enabling effective information transmission and exchange of opinions even in an online environment.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The server retrieves audio streams in real time from the online meeting platform. This includes connecting to the platform via an API and receiving audio data. The audio data is accompanied by metadata such as speaker information.

[0046] Step 2:

[0047] The server inputs the acquired audio stream into a speech recognition engine, which converts it into text format. The speech recognition engine utilizes machine learning models to perform highly accurate text conversion. At this stage, the meeting content is saved as text.

[0048] Step 3:

[0049] The server analyzes the speech-to-text data to understand the context of the conversation. Natural language processing algorithms extract keywords and topics to help understand the content of the conversation. This analysis is then used to generate subsequent responses.

[0050] Step 4:

[0051] The server uses a generation mechanism to generate appropriate responses and interjections based on the analyzed context. This process utilizes a generative AI model, which is required to create responses that naturally match the context and flow of the conversation. For example, responses such as "That's interesting, could you tell me more?" might be generated.

[0052] Step 5:

[0053] The server sends the generated response to the terminal. Since transmission occurs in real time, the communication protocol is optimized to minimize delays.

[0054] Step 6:

[0055] The terminal displays received responses in the user interface. The user interface is designed to allow users to view responses in the form of a chat window or similar. Participants can view these responses and interact as needed.

[0056] Step 7:

[0057] Users offer their opinions and ask questions based on the displayed responses. This interaction makes the conversation more lively and facilitates smoother communication with the speaker. User input is also collected as a log and used to improve the system.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] Achieving smooth, real-time communication in online meetings remains challenging. In voice-based conversations, it's difficult for all participants to understand and react simultaneously, often resulting in one-way dialogue. This can hinder the stimulating nature of discussions and the efficient transmission of information. Therefore, there is a need for a system that instantly transcribes spoken content into text and generates appropriate responses on the spot.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes an acquisition means for collecting audio data via a communication means linked to an electronic device, a conversion means for converting the acquired audio data into text data, and a generation means for understanding the context based on the converted text data and generating an appropriate response. This makes it possible to transcribe speech during a meeting in real time and quickly provide natural responses appropriate to the context.

[0063] "Electronic devices" is a general term for devices used to process or transmit digital information.

[0064] "Communication methods" refer to methods and protocols for sending and receiving information, and include mechanisms for exchanging data over a network.

[0065] "Audio data" refers to information that is recorded or transmitted using digital methods, specifically human speech.

[0066] "Means of acquisition" refers to the process or method for collecting specific information.

[0067] "Text data" refers to information expressed in text format, including audio transcribed into text.

[0068] "Conversion means" refers to technologies and processes for changing data formats or the nature of information.

[0069] "Context" refers to the situation and relationships necessary to understand the meaning and information behind a sentence or statement.

[0070] "Generative means" refers to the technologies and processes used to create new information and data.

[0071] An "information display device" refers to hardware or an interface used to visually present information content such as text and graphics.

[0072] This invention is a system for streamlining real-time communication in online environments. Specifically, it facilitates communication by instantly converting spoken information into text and generating responses during online meetings.

[0073] The server retrieves audio data from the electronic conferencing platform. A common API can be used for this task. For example, when using a web conferencing platform, audio data can be collected in real time through its API.

[0074] The acquired audio data is converted into text data on the server using speech recognition software, such as Google® Speech-to-Text API. This data forms the basis for subsequent analysis and response generation.

[0075] The server utilizes the converted text data to analyze the context using a generative AI model, such as a machine learning algorithm provided by a business partner, and generates an appropriate response. This model is customized to obtain the necessary response using prompt sentences. For example, it might use a prompt sentence like, "Generate appropriate questions for when the speaker is explaining a new product."

[0076] The generated responses are sent to the terminal and displayed on the user interface via an information display device. This interface is designed to be intuitive for users, allowing participants to express their opinions in real time based on the generated responses.

[0077] This allows users to communicate actively even in online meeting environments, thereby improving the quality of meetings.

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] The server acquires audio data in real time from the online meeting platform. The input is the audio stream during the meeting, and the output is digital audio data stored on the server. Specifically, it uses a Web API to continuously retrieve the meeting's audio feed.

[0081] Step 2:

[0082] The server converts the acquired audio data into text data using speech recognition software. The input is the audio data acquired in step 1, and the output is character data in text format. Here, the process of analyzing the audio waveform and converting it into a string of characters in the corresponding language takes place.

[0083] Step 3:

[0084] The server analyzes the context based on the converted text data. The input is the text data generated in step 2, and the server understands its context and extracts relevant keywords. The output is a list of the analyzed contextual information and keywords. Specifically, it uses natural language processing techniques to grasp the intent and focus of the text.

[0085] Step 4:

[0086] The server uses a generative AI model based on contextual information to generate appropriate responses. The input is the contextual information analyzed in step 3, and the output is the generated response text. By providing prompt sentences to the AI ​​model, situation-appropriate questions and responses can be generated. For example, it can create responses such as, "The advantages of the new product are explained, but what other features does it have?"

[0087] Step 5:

[0088] The server sends the generated response to the terminal. The input is the response text generated in step 4, and the output is a chat-style message displayed in the user interface. The terminal visually presents the content to the user and prompts the user to take specific actions to participate in the conversation.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] In online meetings and virtual stores, there is a challenge in generating and presenting efficient and natural responses when participants or customers speak or make inquiries in real time, hindering smooth communication. Furthermore, to improve the quality of information transmission, there is a need for a system that can immediately respond to the diverse needs of participants.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, a generation means for analyzing the context based on the converted speech text and generating an appropriate response, and an additional means for acquiring voice input from a user and generating a corresponding response based on the voice input. This makes it possible to facilitate interaction with participants and customers and improve the quality of information transmission in online meetings and virtual stores.

[0094] "Speech information" refers to audio data acquired during communication using speech, as well as information about the speaker who produced that speech.

[0095] "Acquisition means" refers to technical methods and devices for collecting voice data and necessary information from external platforms or environments.

[0096] A "conversion method" refers to a method or system for converting collected audio data into text information, which can then be used as basic data for understanding the content of a conversation.

[0097] "Generation means" refers to technologies and methods for generating natural responses and comments based on context obtained from speech-to-text.

[0098] "Display means" refers to an interface or system for visually presenting the generated response to the user.

[0099] "Additional means" refers to methods or systems that have the function of receiving direct voice input from the user, processing it, and deriving a response.

[0100] A "virtual store" is a virtual sales environment built on a digital platform as an alternative to a physical store, where users can browse products and make inquiries.

[0101] A "machine learning model" is an algorithm or network that can learn from large amounts of data and generate responses that mimic some aspects of human intelligence.

[0102] The system for realizing this invention is built on the basis of smooth data exchange between servers, terminals, and users.

[0103] The server acquires audio data from online meeting platforms and virtual store environments. Using acquisition methods, the server captures the voice spoken by the user and converts it into text using speech recognition software such as the Google Cloud Speech-to-Text API. This converted text forms the basis for contextual analysis.

[0104] Next, the server uses an OpenAI® generative AI model as a generation mechanism to generate appropriate responses from the speech-to-text format. Based on the prompt, natural responses and questions are generated that understand the flow and context of the conversation. For example, if there is a question about a product, the generative AI model will create a detailed response about its features and specifications.

[0105] The server then sends the generated response to the terminal for display in chat format. The terminal has a user interface through which the user can receive the response and make further inquiries. This user interface operates as an application installed on a smartphone or smart glasses and is developed using technologies such as React Native.

[0106] As a concrete example, if a user asks "Is this product waterproof?" in a virtual store, the server converts the question into text, and the generating AI model produces a response such as "Yes, this product is waterproof."

[0107] As an example of a prompt, the AI ​​model is input in the format, "Generate an appropriate response to the customer from this voice-text data: {voice-text}". Based on this prompt, a response that fits the flow of the conversation is provided, making it possible to facilitate smooth interaction with the user.

[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0109] Step 1:

[0110] The server obtains speech information from the user. This information is obtained via a voice input device (e.g., a smartphone or smart glasses). The input voice data is sent to the server, and this voice data becomes the input.

[0111] Step 2:

[0112] The server converts the acquired audio data into text data using a conversion method. By utilizing the Google Cloud Speech-to-Text API, the audio is converted into text data by converting the speech into a string. The text data is then output.

[0113] Step 3:

[0114] The server uses a generative AI model to analyze the context from text data using a generation method and generate an appropriate response. At this stage, it receives text data as input and inputs prompt sentences into the generative AI model to generate a natural language response that fits the context. The response text is then generated as output.

[0115] Step 4:

[0116] The server sends the generated response to the terminal. The terminal receives this response text and displays it to the user through the user interface. A React Native application is used for display, allowing the user to see the response. A visual representation of the response is provided as output.

[0117] Step 5:

[0118] The user reviews the displayed response through the terminal's interface. They can then enter additional inquiries or comments based on that response. The user's voice or text input becomes new input, and the subsequent dialogue is processed back to step 1.

[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0120] In one embodiment of the present invention, an online communication support system equipped with emotion recognition functionality is provided. This system aims to generate more empathetic and appropriate responses by processing user speech information in real time and recognizing emotions.

[0121] First, the server receives speech information from the online meeting platform. This speech information includes both audio and text data, which are acquired simultaneously. The audio data is converted to text using a speech recognition engine. The converted text is used to understand the content of the conversation.

[0122] Next, the server uses an emotion engine to analyze the user's emotions from the acquired voice and text data. This process infers the emotional state based on the tone, pitch, rhythm, and words used in the voice. For example, if the user is excited, characteristics such as a higher voice tone and faster speech will appear. This emotional information is a crucial element in making responses more human and empathetic.

[0123] Based on the emotional information obtained by the emotion engine, the server uses generation methods to produce contextually appropriate responses. The generated responses reflect the emotional state and take into account the corresponding tone and content. For example, if a user expresses anxiety, the system will generate a reassuring response such as, "Are you okay? Is there anything I can do to help?"

[0124] The generated response is sent from the server to the terminal and displayed on the user interface. The display on the terminal is designed to be intuitive and easy for the user to understand. Based on this, the user can participate in the conversation and engage in active real-time communication.

[0125] This system enables deeper mutual understanding among participants in online communication through responses that take user emotions into account. As a result, it is expected to significantly improve the efficiency and satisfaction of meetings and lectures.

[0126] The following describes the processing flow.

[0127] Step 1:

[0128] The server retrieves audio streams and metadata in real time from the online meeting platform. The audio streams, which include the speech of all participants during the meeting, are sent to the server via the platform's API.

[0129] Step 2:

[0130] The server processes the acquired audio stream through a speech recognition engine to convert it to text. This speech recognition engine uses machine learning algorithms to achieve highly accurate transcription and quickly generates text from the acquired audio data.

[0131] Step 3:

[0132] The emotion engine built into the server analyzes both audio and text data to identify the speaker's emotional state. It detects voice tone, speed, and keywords, and infers emotions (such as joy, anger, sadness, etc.) from this information. For example, a high-pitched and fast voice may indicate tension or excitement.

[0133] Step 4:

[0134] The server automatically generates an appropriate response using a generation mechanism based on the evaluation results of the emotional state and the voice-to-text. This response takes into account the user's current emotions and is designed to ensure a smooth flow of conversation. For example, if anxiety is detected, it might respond with something like, "Please calm down, and let me know if there's anything bothering you."

[0135] Step 5:

[0136] The server sends the generated response to the terminal and displays it on the user interface. The display method uses a visually simple and easy-to-understand format so that the user can quickly comprehend it.

[0137] Step 6:

[0138] Users review the responses displayed on their devices and input their own opinions and questions. Further responses and new interactions are then initiated, taking user feedback into consideration. This allows the meeting conversation to progress naturally and deepens understanding among participants.

[0139] This system facilitates smoother communication in online meetings and lectures by providing responses tailored to the user's emotional state. As a result, it is expected to improve participant satisfaction and enhance the effectiveness of the meetings.

[0140] (Example 2)

[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0142] Online communication presents a challenge in instantly grasping the subtle nuances of participants' emotions and providing appropriate responses. This can lead to misunderstandings and inefficiencies in meetings and discussions. To address this, a system is needed that analyzes emotions in real time and generates empathetic responses that correspond to the user's feelings.

[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0144] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into language data, and a generation means for analyzing emotions based on the converted language data and phonological features and generating a contextually appropriate response. This enables real-time, empathetic response generation that takes the user's emotions into consideration.

[0145] "Speech information" refers to conversational data during communication, including speech and associated text data.

[0146] "Acquisition means" refers to a function or method for collecting speech information from an external source and incorporating it into the system.

[0147] "Conversion means" refers to a function or method for converting acquired audio data into text data, i.e., language data.

[0148] "Language data" refers to written text information converted from audio data using speech recognition technology.

[0149] "Phonological features" refer to information that characterizes the structure of speech, including elements such as tone, pitch, and rhythm.

[0150] "Emotional analysis" is the process of using voice and language data to infer and classify the user's emotional state.

[0151] "Generating means" refers to a function or method for producing an appropriate response based on information obtained through analysis.

[0152] "Contextually appropriate responses" refer to natural and empathetic responses that are based on the user's statements and emotional state.

[0153] "Display means" refers to a function or method for visually presenting the generated response on a user interface.

[0154] This invention relates to an online communication support system equipped with emotion recognition capabilities. The system aims to collect user speech information in real time, analyze emotions, and generate appropriate responses.

[0155] The server acquires speech information from the online meeting platform. This information includes audio and text data, which are incorporated into the system through the acquisition mechanism. Using speech recognition software (e.g., a speech recognition API), the server converts the audio data into text data. The converted text data serves as the basis for analyzing the content of the conversation.

[0156] Next, the server uses emotion pattern recognition software (e.g., emotion analysis engine) to analyze the phonological features contained in the audio data. This analysis infers and classifies the speaker's emotional state. Based on this information, the server utilizes a generative AI model (e.g., natural language generation model) to generate a contextually appropriate response.

[0157] The generated response is sent to the terminal via the network. The terminal displays the response in its user interface, allowing the user to understand it immediately and facilitate smooth communication.

[0158] For example, if a user says, "I'm worried because the project is behind schedule," the server performs sentiment analysis on that utterance and assigns a label indicating anxiety. Next, it uses a generative AI model to generate a reassuring response such as, "Let's think together about how we can support you."

[0159] Examples of prompts to input into a generative AI model:

[0160] User utterance: "I'm worried because the project is behind schedule."

[0161] Emotional labeling based on analysis: Anxiety

[0162] Response generation prompt: "What kind of reassurance can we provide to anxious users?"

[0163] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0164] Step 1:

[0165] The server acquires speech information in real time from the online meeting platform. The input consists of audio data and text data, and it receives both types of data and begins processing them simultaneously. The audio data is sent to the speech recognition engine for processing, and the text data is sent to the sentiment analysis engine.

[0166] Step 2:

[0167] The server uses speech recognition software to convert the received audio data into language data. Specifically, it converts the audio file into a byte stream format and sends it to the speech recognition API. The input is audio data, and the output is string data that reflects the content of the audio. This text is used in subsequent parsing steps.

[0168] Step 3:

[0169] The server analyzes emotions using an emotion analysis engine based on linguistic data and phonological features. It analyzes voice tone, pitch, and rhythm with high accuracy and performs semantic analysis on this data along with text data. The input consists of phonological feature parameters and text data, and the output consists of emotion labels (e.g., joy, sadness, anger) and numerical data indicating their intensity.

[0170] Step 4:

[0171] Using a generative AI model, appropriate responses are generated using analyzed sentiment labels. The server prepares sentiment labels and text data as prompts and inputs them into the AI ​​model. Based on this input, the AI ​​model generates a contextually appropriate and natural response. The input consists of sentiment labels and prompts, and the output is text data to be sent back to the user.

[0172] Step 5:

[0173] The server sends the generated response to the terminal. The terminal displays the received response text in its user interface. Specifically, it creates a notification on the screen based on the response text and presents it to the user in an easily understandable format, either audibly or visually. The input is the generated response text, and the output is the visible display shown on the screen.

[0174] (Application Example 2)

[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0176] In autonomous vehicles, there is a challenge in providing an appropriate in-vehicle environment that responds to the emotional state of passengers. If it were possible to respond to passengers' emotions, travel in autonomous vehicles would be more comfortable and safer. Conventional technologies have been unable to accurately grasp passengers' emotions and provide responses based on them in real time.

[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0178] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, an analysis means for analyzing emotions based on the converted speech text and speech features, a generation means for generating an appropriate response based on the emotion information obtained by the analysis, and a display means for displaying the generated response on a user input device. This enables real-time responses and adjustments to the in-vehicle environment in accordance with the emotional state of passengers.

[0179] "Speech information" refers to audio waves and associated text data, which are data acquired during communication.

[0180] "Acquisition means" refers to a device or method for receiving underlying data from voice input devices or sensors.

[0181] "Speech-to-text" refers to textual information generated from sound waves using speech recognition technology.

[0182] "Conversion means" refers to technology or equipment for converting sound waves into textual information.

[0183] "Means of analyzing emotions" refers to technologies or devices for inferring the emotional state of a speaker from voice and text data.

[0184] "Generation means" refers to a method or apparatus for creating appropriate responses or actions based on analyzed data.

[0185] "Display means" refers to a method or apparatus for providing a generated response to a user visually or audibly.

[0186] A "user input device" refers to a device used by a user to receive or manipulate information.

[0187] The system for carrying out this invention comprises several technical elements. The server collects sound wave data from microphones installed in the vehicle using an acquisition means for acquiring passenger speech information. This sound wave data is then converted into speech-to-text by a conversion means and processed as text data.

[0188] The server uses analytical tools responsible for sentiment analysis to evaluate the passenger's emotional state from the characteristics of the speech text and sound waves. This analysis takes into account factors such as the tone, tempo, and intensity of the speech, as well as any bias in the vocabulary used. The results of the analysis are recorded in the server as the passenger's emotional state.

[0189] Next, the server uses a generation mechanism to generate a response based on the acquired emotional information. This response includes in-car voice announcements that take the passenger's emotions into consideration, automatically selected background music, and voice guidance. For example, if it is determined that a passenger is feeling tense, relaxing background music may be automatically played.

[0190] The responses displayed on the terminal are provided by a display mechanism and presented in a format that is easy for passengers to understand and use. Through this series of technological elements, passenger emotional satisfaction is enhanced, enabling a safe and comfortable autonomous vehicle experience.

[0191] For example, if the conversation in the car during a family trip is cheerful and positive, the system will select upbeat background music and take measures to enhance the atmosphere inside the car. Another example of a prompt to be input to the generation AI model is in the format of "Detect the emotion of the following sentence and generate the corresponding feedback: Text." This allows the system to identify the passenger's emotion from the text and provide the most appropriate in-car response.

[0192] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0193] Step 1:

[0194] The server acquires passenger speech information from microphones installed inside the vehicle. The input is audio wave data. The server collects this audio wave data using an acquisition method and stores it in a format that can be processed in the next step.

[0195] Step 2:

[0196] The server converts audio wave data into speech-to-text using a conversion mechanism. The input is audio wave data, and the output is speech-to-text. In this process, speech recognition software analyzes the characteristics of the sound waves and generates corresponding character information. Specifically, it analyzes the amplitude and frequency of the sound waves and converts them into Japanese words and phrases.

[0197] Step 3:

[0198] The server analyzes the characteristics of the generated speech-text and speech-wave data and uses analytical means to estimate the passenger's emotions. The input is speech-text and speech-wave features, and the output is emotion data. This emotion analysis takes into account voice tone, speed, and vocabulary used to estimate a specific emotional state (e.g., joy, sadness, tension).

[0199] Step 4:

[0200] The server utilizes generation methods to generate appropriate responses based on the passenger's emotional state, using emotion analysis results. The input is emotion data, and the output is response text and associated in-vehicle operation commands. This step generates text for emotion-appropriate background music playback and voice guidance. A generation AI model is used to design natural, situation-aware responses.

[0201] Step 5:

[0202] The server sends the generated response to the terminal and displays it on the in-vehicle user interface. Inputs are response text and in-vehicle operation commands, and outputs are the in-vehicle responses perceived by the user (music playback, voice guidance, etc.). The terminal provides passengers with specific responses via voice, enabling them to take action or make choices based on them.

[0203] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0204] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0205] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0206] [Second Embodiment]

[0207] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0208] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0209] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0210] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0211] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0212] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0213] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0214] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0215] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0216] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0217] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0218] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0219] In one embodiment of the invention, a system is constructed to improve real-time communication in online meetings and lectures. The specific operation of this system is described below.

[0220] First, the server acquires speech information in real time from the online meeting platform. This information includes voice data and speaker IDs, and is continuously collected as the meeting progresses.

[0221] Next, the server converts the acquired audio data into text format using speech recognition technology. This converted audio-text becomes the basis for analyzing the content of the conversation. For example, if the speaker is discussing a new product, relevant keywords and sentences will be reflected in the text.

[0222] Next, the server's generation system analyzes the context using the generated text and produces appropriate responses and comments. This generation process utilizes machine learning models to generate natural interjections and questions that fit the flow of the conversation. For example, if the speaker is talking about a new feature, the AI ​​might create a response such as, "How will that feature improve our work?"

[0223] The generated response is immediately sent from the server to the terminal. On the terminal, this response is displayed in a chat format via a user interface, allowing participants to review it and enter their own opinions or questions as needed. This user interface is designed with ease of use and visibility in mind, making it intuitive for participants to use.

[0224] Furthermore, by allowing users to participate in the discussion using interjections and comments provided through this system, speakers can instantly grasp the audience's interest and level of understanding. This enables speakers to adjust the information they provide in real time, which is expected to improve the quality of online meetings.

[0225] By utilizing such a system, smooth communication can be promoted for both speakers and listeners, enabling effective information transmission and exchange of opinions even in an online environment.

[0226] The following describes the processing flow.

[0227] Step 1:

[0228] The server retrieves audio streams in real time from the online meeting platform. This includes connecting to the platform via an API and receiving audio data. The audio data is accompanied by metadata such as speaker information.

[0229] Step 2:

[0230] The server inputs the acquired audio stream into a speech recognition engine, which converts it into text format. The speech recognition engine utilizes machine learning models to perform highly accurate text conversion. At this stage, the meeting content is saved as text.

[0231] Step 3:

[0232] The server analyzes the speech-to-text data to understand the context of the conversation. Natural language processing algorithms extract keywords and topics to help understand the content of the conversation. This analysis is then used to generate subsequent responses.

[0233] Step 4:

[0234] The server uses a generation mechanism to generate appropriate responses and interjections based on the analyzed context. This process utilizes a generative AI model, which is required to create responses that naturally match the context and flow of the conversation. For example, responses such as "That's interesting, could you tell me more?" might be generated.

[0235] Step 5:

[0236] The server sends the generated response to the terminal. Since transmission occurs in real time, the communication protocol is optimized to minimize delays.

[0237] Step 6:

[0238] The terminal displays received responses in the user interface. The user interface is designed to allow users to view responses in the form of a chat window or similar. Participants can view these responses and interact as needed.

[0239] Step 7:

[0240] Users offer their opinions and ask questions based on the displayed responses. This interaction makes the conversation more lively and facilitates smoother communication with the speaker. User input is also collected as a log and used to improve the system.

[0241] (Example 1)

[0242] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0243] Achieving smooth, real-time communication in online meetings remains challenging. In voice-based conversations, it's difficult for all participants to understand and react simultaneously, often resulting in one-way dialogue. This can hinder the stimulating nature of discussions and the efficient transmission of information. Therefore, there is a need for a system that instantly transcribes spoken content into text and generates appropriate responses on the spot.

[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0245] In this invention, the server includes an acquisition means for collecting audio data via a communication means linked to an electronic device, a conversion means for converting the acquired audio data into text data, and a generation means for understanding the context based on the converted text data and generating an appropriate response. This makes it possible to transcribe speech during a meeting in real time and quickly provide natural responses appropriate to the context.

[0246] "Electronic devices" is a general term for devices used to process or transmit digital information.

[0247] "Communication methods" refer to methods and protocols for sending and receiving information, and include mechanisms for exchanging data over a network.

[0248] "Audio data" refers to information that is recorded or transmitted using digital methods, specifically human speech.

[0249] "Means of acquisition" refers to the process or method for collecting specific information.

[0250] "Text data" refers to information expressed in text format, including audio transcribed into text.

[0251] "Conversion means" refers to technologies and processes for changing data formats or the nature of information.

[0252] "Context" refers to the situation and relationships necessary to understand the meaning and information behind a sentence or statement.

[0253] "Generative means" refers to the technologies and processes used to create new information and data.

[0254] An "information display device" refers to hardware or an interface used to visually present information content such as text and graphics.

[0255] This invention is a system for streamlining real-time communication in online environments. Specifically, it facilitates communication by instantly converting spoken information into text and generating responses during online meetings.

[0256] The server retrieves audio data from the electronic conferencing platform. A common API can be used for this task. For example, when using a web conferencing platform, audio data can be collected in real time through its API.

[0257] The acquired audio data is converted into text data on the server using speech recognition software, such as the Google Speech-to-Text API. This data forms the basis for subsequent analysis and response generation.

[0258] The server utilizes the converted text data to analyze the context using a generative AI model, such as a machine learning algorithm provided by a business partner, and generates an appropriate response. This model is customized to obtain the necessary response using prompt sentences. For example, it might use a prompt sentence like, "Generate appropriate questions for when the speaker is explaining a new product."

[0259] The generated responses are sent to the terminal and displayed on the user interface via an information display device. This interface is designed to be intuitive for users, allowing participants to express their opinions in real time based on the generated responses.

[0260] This allows users to communicate actively even in online meeting environments, thereby improving the quality of meetings.

[0261] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0262] Step 1:

[0263] The server acquires audio data in real time from the online meeting platform. The input is the audio stream during the meeting, and the output is digital audio data stored on the server. Specifically, it uses a Web API to continuously retrieve the meeting's audio feed.

[0264] Step 2:

[0265] The server converts the acquired audio data into text data using speech recognition software. The input is the audio data acquired in step 1, and the output is character data in text format. Here, the process of analyzing the audio waveform and converting it into a string of characters in the corresponding language takes place.

[0266] Step 3:

[0267] The server analyzes the context based on the converted text data. The input is the text data generated in step 2, and the server understands its context and extracts relevant keywords. The output is a list of the analyzed contextual information and keywords. Specifically, it uses natural language processing techniques to grasp the intent and focus of the text.

[0268] Step 4:

[0269] The server uses a generative AI model based on contextual information to generate appropriate responses. The input is the contextual information analyzed in step 3, and the output is the generated response text. By providing prompt sentences to the AI ​​model, situation-appropriate questions and responses can be generated. For example, it can create responses such as, "The advantages of the new product are explained, but what other features does it have?"

[0270] Step 5:

[0271] The server sends the generated response to the terminal. The input is the response text generated in step 4, and the output is a chat-style message displayed in the user interface. The terminal visually presents the content to the user and prompts the user to take specific actions to participate in the conversation.

[0272] (Application Example 1)

[0273] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0274] In online meetings and virtual stores, there is a challenge in generating and presenting efficient and natural responses when participants or customers speak or make inquiries in real time, hindering smooth communication. Furthermore, to improve the quality of information transmission, there is a need for a system that can immediately respond to the diverse needs of participants.

[0275] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0276] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, a generation means for analyzing the context based on the converted speech text and generating an appropriate response, and an additional means for acquiring voice input from a user and generating a corresponding response based on the voice input. This makes it possible to facilitate interaction with participants and customers and improve the quality of information transmission in online meetings and virtual stores.

[0277] "Speech information" refers to audio data acquired during communication using speech, as well as information about the speaker who produced that speech.

[0278] "Acquisition means" refers to technical methods and devices for collecting voice data and necessary information from external platforms or environments.

[0279] A "conversion method" refers to a method or system for converting collected audio data into text information, which can then be used as basic data for understanding the content of a conversation.

[0280] "Generation means" refers to technologies and methods for generating natural responses and comments based on context obtained from speech-to-text.

[0281] The "display means" is an interface or system for visually presenting the generated response to the user.

[0282] The "additional means" is a method or system that has the function of receiving direct voice input from the user, processing it, and deriving a response.

[0283] A "virtual store" is a virtual sales environment constructed on a digital platform as an alternative to a physical store, where users can view products and make inquiries.

[0284] A "machine learning model" is an algorithm or network that can learn from a large amount of data and generate responses by imitating a part of human intelligence.

[0285] The system for realizing this invention is constructed based on the smooth data exchange among the server, the terminal, and the user.

[0286] The server acquires voice data from an online meeting platform or a virtual store environment. Here, the server uses acquisition means to acquire the voice uttered by the user and converts it into voice text using voice recognition software such as the Google Cloud Speech-to-Text API. This converted text serves as the basis for analyzing the context.

[0287] Next, as generation means, the server uses the generation AI model of OpenAI to generate an appropriate response from the voice text. Based on the prompt, natural responses and questions that understand the flow and context of the conversation are generated. For example, when there is a question about a product, the generation AI model creates a detailed response about its features and specifications.

[0288] The server then sends the generated response to the terminal for display in chat format. The terminal has a user interface through which the user can receive the response and make further inquiries. This user interface operates as an application installed on a smartphone or smart glasses and is developed using technologies such as React Native.

[0289] As a concrete example, if a user asks "Is this product waterproof?" in a virtual store, the server converts the question into text, and the generating AI model produces a response such as "Yes, this product is waterproof."

[0290] As an example of a prompt, the AI ​​model is input in the format, "Generate an appropriate response to the customer from this voice-text data: {voice-text}". Based on this prompt, a response that fits the flow of the conversation is provided, making it possible to facilitate smooth interaction with the user.

[0291] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0292] Step 1:

[0293] The server obtains speech information from the user. This information is obtained via a voice input device (e.g., a smartphone or smart glasses). The input voice data is sent to the server, and this voice data becomes the input.

[0294] Step 2:

[0295] The server converts the acquired audio data into text data using a conversion method. By utilizing the Google Cloud Speech-to-Text API, the audio is converted into text data by converting the speech into a string. The text data is then output.

[0296] Step 3:

[0297] The server uses a generative AI model to analyze the context from text data using a generation method and generate an appropriate response. At this stage, it receives text data as input and inputs prompt sentences into the generative AI model to generate a natural language response that fits the context. The response text is then generated as output.

[0298] Step 4:

[0299] The server sends the generated response to the terminal. The terminal receives this response text and displays it to the user through the user interface. A React Native application is used for display, allowing the user to see the response. A visual representation of the response is provided as output.

[0300] Step 5:

[0301] The user reviews the displayed response through the terminal's interface. They can then enter additional inquiries or comments based on that response. The user's voice or text input becomes new input, and the subsequent dialogue is processed back to step 1.

[0302] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0303] In one embodiment of the present invention, an online communication support system equipped with emotion recognition functionality is provided. This system aims to generate more empathetic and appropriate responses by processing user speech information in real time and recognizing emotions.

[0304] First, the server receives speech information from the online meeting platform. This speech information includes audio data and text data, which are acquired simultaneously. The audio data is converted into speech text using a speech recognition engine. The converted text is used to understand the content of the conversation.

[0305] Next, the server utilizes an emotion engine to analyze the user's emotion from the acquired audio and text data. In this process, the emotional state is inferred based on the tone, pitch, rhythm of the voice, and the words used. For example, when the user is excited, characteristics such as an increased tone of voice and faster speech patterns appear. This emotion information is an important factor for making responses more human and empathetic.

[0306] Based on the emotion information obtained by the emotion engine, the server utilizes a generation means to generate a response suitable for the context. The generated response reflects the emotional state and takes into account the corresponding tone and content. For example, when the user shows anxiety, the system generates a response that gives a sense of reassurance such as "Are you okay? Is there anything I can support you with?"

[0307] The generated response is sent from the server to the terminal and displayed on the user interface. The display on the terminal is designed to be in a form that allows the user to intuitively understand the information easily. Based on this, the user can participate in the conversation and actively conduct real-time communication.

[0308] This system enables deeper mutual understanding among participants in online communication through responses that take into account the user's emotions. As a result, it is expected to significantly improve the efficiency and satisfaction of meetings and lectures.

[0309] The following describes the processing flow.

[0310] Step 1:

[0311] The server retrieves audio streams and metadata in real time from the online meeting platform. The audio streams, which include the speech of all participants during the meeting, are sent to the server via the platform's API.

[0312] Step 2:

[0313] The server processes the acquired audio stream through a speech recognition engine to convert it to text. This speech recognition engine uses machine learning algorithms to achieve highly accurate transcription and quickly generates text from the acquired audio data.

[0314] Step 3:

[0315] The emotion engine built into the server analyzes both audio and text data to identify the speaker's emotional state. It detects voice tone, speed, and keywords, and infers emotions (such as joy, anger, sadness, etc.) from this information. For example, a high-pitched and fast voice may indicate tension or excitement.

[0316] Step 4:

[0317] The server automatically generates an appropriate response using a generation mechanism based on the evaluation results of the emotional state and the voice-to-text. This response takes into account the user's current emotions and is designed to ensure a smooth flow of conversation. For example, if anxiety is detected, it might respond with something like, "Please calm down, and let me know if there's anything bothering you."

[0318] Step 5:

[0319] The server sends the generated response to the terminal and displays it on the user interface. The display method uses a visually simple and easy-to-understand format so that the user can quickly comprehend it.

[0320] Step 6:

[0321] Users review the responses displayed on their devices and input their own opinions and questions. Further responses and new interactions are then initiated, taking user feedback into consideration. This allows the meeting conversation to progress naturally and deepens understanding among participants.

[0322] This system facilitates smoother communication in online meetings and lectures by providing responses tailored to the user's emotional state. As a result, it is expected to improve participant satisfaction and enhance the effectiveness of the meetings.

[0323] (Example 2)

[0324] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0325] Online communication presents a challenge in instantly grasping the subtle nuances of participants' emotions and providing appropriate responses. This can lead to misunderstandings and inefficiencies in meetings and discussions. To address this, a system is needed that analyzes emotions in real time and generates empathetic responses that correspond to the user's feelings.

[0326] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0327] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into language data, and a generation means for analyzing emotions based on the converted language data and phonological features and generating a contextually appropriate response. This enables real-time, empathetic response generation that takes the user's emotions into consideration.

[0328] "Speech information" refers to conversational data during communication, including speech and associated text data.

[0329] "Acquisition means" refers to a function or method for collecting speech information from an external source and incorporating it into the system.

[0330] "Conversion means" refers to a function or method for converting acquired audio data into text data, i.e., language data.

[0331] "Language data" refers to written text information converted from audio data using speech recognition technology.

[0332] "Phonological features" refer to information that characterizes the structure of speech, including elements such as tone, pitch, and rhythm.

[0333] "Emotional analysis" is the process of using voice and language data to infer and classify the user's emotional state.

[0334] "Generating means" refers to a function or method for producing an appropriate response based on information obtained through analysis.

[0335] "Contextually appropriate responses" refer to natural and empathetic responses that are based on the user's statements and emotional state.

[0336] "Display means" refers to a function or method for visually presenting the generated response on a user interface.

[0337] This invention relates to an online communication support system equipped with emotion recognition capabilities. The system aims to collect user speech information in real time, analyze emotions, and generate appropriate responses.

[0338] The server acquires speech information from the online meeting platform. This information includes audio and text data, which are incorporated into the system through the acquisition mechanism. Using speech recognition software (e.g., a speech recognition API), the server converts the audio data into text data. The converted text data serves as the basis for analyzing the content of the conversation.

[0339] Next, the server uses emotion pattern recognition software (e.g., emotion analysis engine) to analyze the phonological features contained in the audio data. This analysis infers and classifies the speaker's emotional state. Based on this information, the server utilizes a generative AI model (e.g., natural language generation model) to generate a contextually appropriate response.

[0340] The generated response is sent to the terminal via the network. The terminal displays the response in its user interface, allowing the user to understand it immediately and facilitate smooth communication.

[0341] For example, if a user says, "I'm worried because the project is behind schedule," the server performs sentiment analysis on that utterance and assigns a label indicating anxiety. Next, it uses a generative AI model to generate a reassuring response such as, "Let's think together about how we can support you."

[0342] Examples of prompts to input into a generative AI model:

[0343] User utterance: "I'm worried because the project is behind schedule."

[0344] Emotional labeling based on analysis: Anxiety

[0345] Response generation prompt: "What kind of reassurance can we provide to anxious users?"

[0346] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0347] Step 1:

[0348] The server acquires speech information in real time from the online meeting platform. The input consists of audio data and text data, and it receives both types of data and begins processing them simultaneously. The audio data is sent to the speech recognition engine for processing, and the text data is sent to the sentiment analysis engine.

[0349] Step 2:

[0350] The server uses speech recognition software to convert the received audio data into language data. Specifically, it converts the audio file into a byte stream format and sends it to the speech recognition API. The input is audio data, and the output is string data that reflects the content of the audio. This text is used in subsequent parsing steps.

[0351] Step 3:

[0352] The server analyzes emotions using an emotion analysis engine based on linguistic data and phonological features. It analyzes voice tone, pitch, and rhythm with high accuracy and performs semantic analysis on this data along with text data. The input consists of phonological feature parameters and text data, and the output consists of emotion labels (e.g., joy, sadness, anger) and numerical data indicating their intensity.

[0353] Step 4:

[0354] Using a generative AI model, appropriate responses are generated using analyzed sentiment labels. The server prepares sentiment labels and text data as prompts and inputs them into the AI ​​model. Based on this input, the AI ​​model generates a contextually appropriate and natural response. The input consists of sentiment labels and prompts, and the output is text data to be sent back to the user.

[0355] Step 5:

[0356] The server sends the generated response to the terminal. The terminal displays the received response text in its user interface. Specifically, it creates a notification on the screen based on the response text and presents it to the user in an easily understandable format, either audibly or visually. The input is the generated response text, and the output is the visible display shown on the screen.

[0357] (Application Example 2)

[0358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0359] In autonomous vehicles, there is a challenge in providing an appropriate in-vehicle environment that responds to the emotional state of passengers. If it were possible to respond to passengers' emotions, travel in autonomous vehicles would be more comfortable and safer. Conventional technologies have been unable to accurately grasp passengers' emotions and provide responses based on them in real time.

[0360] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0361] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, an analysis means for analyzing emotions based on the converted speech text and speech features, a generation means for generating an appropriate response based on the emotion information obtained by the analysis, and a display means for displaying the generated response on a user input device. This enables real-time responses and adjustments to the in-vehicle environment in accordance with the emotional state of passengers.

[0362] "Speech information" refers to audio waves and associated text data, which are data acquired during communication.

[0363] "Acquisition means" refers to a device or method for receiving underlying data from voice input devices or sensors.

[0364] "Speech-to-text" refers to textual information generated from sound waves using speech recognition technology.

[0365] "Conversion means" refers to technology or equipment for converting sound waves into textual information.

[0366] "Means of analyzing emotions" refers to technologies or devices for inferring the emotional state of a speaker from voice and text data.

[0367] "Generation means" refers to a method or apparatus for creating appropriate responses or actions based on analyzed data.

[0368] "Display means" refers to a method or apparatus for providing a generated response to a user visually or audibly.

[0369] A "user input device" refers to a device used by a user to receive or manipulate information.

[0370] The system for carrying out this invention comprises several technical elements. The server collects sound wave data from microphones installed in the vehicle using an acquisition means for acquiring passenger speech information. This sound wave data is then converted into speech-to-text by a conversion means and processed as text data.

[0371] The server uses analytical tools responsible for sentiment analysis to evaluate the passenger's emotional state from the characteristics of the speech text and sound waves. This analysis takes into account factors such as the tone, tempo, and intensity of the speech, as well as any bias in the vocabulary used. The results of the analysis are recorded in the server as the passenger's emotional state.

[0372] Next, the server uses a generation mechanism to generate a response based on the acquired emotional information. This response includes in-car voice announcements that take the passenger's emotions into consideration, automatically selected background music, and voice guidance. For example, if it is determined that a passenger is feeling tense, relaxing background music may be automatically played.

[0373] The responses displayed on the terminal are provided by a display mechanism and presented in a format that is easy for passengers to understand and use. Through this series of technological elements, passenger emotional satisfaction is enhanced, enabling a safe and comfortable autonomous vehicle experience.

[0374] For example, if the conversation in the car during a family trip is cheerful and positive, the system will select upbeat background music and take measures to enhance the atmosphere inside the car. Another example of a prompt to be input to the generation AI model is in the format of "Detect the emotion of the following sentence and generate the corresponding feedback: Text." This allows the system to identify the passenger's emotion from the text and provide the most appropriate in-car response.

[0375] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0376] Step 1:

[0377] The server acquires passenger speech information from microphones installed inside the vehicle. The input is audio wave data. The server collects this audio wave data using an acquisition method and stores it in a format that can be processed in the next step.

[0378] Step 2:

[0379] The server converts audio wave data into speech-to-text using a conversion mechanism. The input is audio wave data, and the output is speech-to-text. In this process, speech recognition software analyzes the characteristics of the sound waves and generates corresponding character information. Specifically, it analyzes the amplitude and frequency of the sound waves and converts them into Japanese words and phrases.

[0380] Step 3:

[0381] The server analyzes the characteristics of the generated speech-text and speech-wave data and uses analytical means to estimate the passenger's emotions. The input is speech-text and speech-wave features, and the output is emotion data. This emotion analysis takes into account voice tone, speed, and vocabulary used to estimate a specific emotional state (e.g., joy, sadness, tension).

[0382] Step 4:

[0383] The server utilizes generation methods to generate appropriate responses based on the passenger's emotional state, using emotion analysis results. The input is emotion data, and the output is response text and associated in-vehicle operation commands. This step generates text for emotion-appropriate background music playback and voice guidance. A generation AI model is used to design natural, situation-aware responses.

[0384] Step 5:

[0385] The server sends the generated response to the terminal and displays it on the in-vehicle user interface. Inputs are response text and in-vehicle operation commands, and outputs are the in-vehicle responses perceived by the user (music playback, voice guidance, etc.). The terminal provides passengers with specific responses via voice, enabling them to take action or make choices based on them.

[0386] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0387] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0388] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0389] [Third Embodiment]

[0390] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0391] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0392] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0393] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0394] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0395] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0396] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0397] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0398] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0399] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0400] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0401] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0402] In one embodiment of the invention, a system is constructed to improve real-time communication in online meetings and lectures. The specific operation of this system is described below.

[0403] First, the server acquires speech information in real time from the online meeting platform. This information includes voice data and speaker IDs, and is continuously collected as the meeting progresses.

[0404] Next, the server converts the acquired audio data into text format using speech recognition technology. This converted audio-text becomes the basis for analyzing the content of the conversation. For example, if the speaker is discussing a new product, relevant keywords and sentences will be reflected in the text.

[0405] Next, the server's generation system analyzes the context using the generated text and produces appropriate responses and comments. This generation process utilizes machine learning models to generate natural interjections and questions that fit the flow of the conversation. For example, if the speaker is talking about a new feature, the AI ​​might create a response such as, "How will that feature improve our work?"

[0406] The generated response is immediately sent from the server to the terminal. On the terminal, this response is displayed in a chat format via a user interface, allowing participants to review it and enter their own opinions or questions as needed. This user interface is designed with ease of use and visibility in mind, making it intuitive for participants to use.

[0407] Furthermore, by allowing users to participate in the discussion using interjections and comments provided through this system, speakers can instantly grasp the audience's interest and level of understanding. This enables speakers to adjust the information they provide in real time, which is expected to improve the quality of online meetings.

[0408] By utilizing such a system, smooth communication can be promoted for both speakers and listeners, enabling effective information transmission and exchange of opinions even in an online environment.

[0409] The following describes the processing flow.

[0410] Step 1:

[0411] The server retrieves audio streams in real time from the online meeting platform. This includes connecting to the platform via an API and receiving audio data. The audio data is accompanied by metadata such as speaker information.

[0412] Step 2:

[0413] The server inputs the acquired audio stream into a speech recognition engine, which converts it into text format. The speech recognition engine utilizes machine learning models to perform highly accurate text conversion. At this stage, the meeting content is saved as text.

[0414] Step 3:

[0415] The server analyzes the speech-to-text data to understand the context of the conversation. Natural language processing algorithms extract keywords and topics to help understand the content of the conversation. This analysis is then used to generate subsequent responses.

[0416] Step 4:

[0417] The server uses a generation mechanism to generate appropriate responses and interjections based on the analyzed context. This process utilizes a generative AI model, which is required to create responses that naturally match the context and flow of the conversation. For example, responses such as "That's interesting, could you tell me more?" might be generated.

[0418] Step 5:

[0419] The server sends the generated response to the terminal. Since transmission occurs in real time, the communication protocol is optimized to minimize delays.

[0420] Step 6:

[0421] The terminal displays received responses in the user interface. The user interface is designed to allow users to view responses in the form of a chat window or similar. Participants can view these responses and interact as needed.

[0422] Step 7:

[0423] Users offer their opinions and ask questions based on the displayed responses. This interaction makes the conversation more lively and facilitates smoother communication with the speaker. User input is also collected as a log and used to improve the system.

[0424] (Example 1)

[0425] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0426] Achieving smooth, real-time communication in online meetings remains challenging. In voice-based conversations, it's difficult for all participants to understand and react simultaneously, often resulting in one-way dialogue. This can hinder the stimulating nature of discussions and the efficient transmission of information. Therefore, there is a need for a system that instantly transcribes spoken content into text and generates appropriate responses on the spot.

[0427] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0428] In this invention, the server includes an acquisition means for collecting audio data via a communication means linked to an electronic device, a conversion means for converting the acquired audio data into text data, and a generation means for understanding the context based on the converted text data and generating an appropriate response. This makes it possible to transcribe speech during a meeting in real time and quickly provide natural responses appropriate to the context.

[0429] "Electronic devices" is a general term for devices used to process or transmit digital information.

[0430] "Communication methods" refer to methods and protocols for sending and receiving information, and include mechanisms for exchanging data over a network.

[0431] "Audio data" refers to information that is recorded or transmitted using digital methods, specifically human speech.

[0432] "Means of acquisition" refers to the process or method for collecting specific information.

[0433] "Text data" refers to information expressed in text format, including audio transcribed into text.

[0434] "Conversion means" refers to technologies and processes for changing data formats or the nature of information.

[0435] "Context" refers to the situation and relationships necessary to understand the meaning and information behind a sentence or statement.

[0436] "Generative means" refers to the technologies and processes used to create new information and data.

[0437] An "information display device" refers to hardware or an interface used to visually present information content such as text and graphics.

[0438] This invention is a system for streamlining real-time communication in online environments. Specifically, it facilitates communication by instantly converting spoken information into text and generating responses during online meetings.

[0439] The server retrieves audio data from the electronic conferencing platform. A common API can be used for this task. For example, when using a web conferencing platform, audio data can be collected in real time through its API.

[0440] The acquired audio data is converted into text data on the server using speech recognition software, such as the Google Speech-to-Text API. This data forms the basis for subsequent analysis and response generation.

[0441] The server utilizes the converted text data to analyze the context using a generative AI model, such as a machine learning algorithm provided by a business partner, and generates an appropriate response. This model is customized to obtain the necessary response using prompt sentences. For example, it might use a prompt sentence like, "Generate appropriate questions for when the speaker is explaining a new product."

[0442] The generated responses are sent to the terminal and displayed on the user interface via an information display device. This interface is designed to be intuitive for users, allowing participants to express their opinions in real time based on the generated responses.

[0443] This allows users to communicate actively even in online meeting environments, thereby improving the quality of meetings.

[0444] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0445] Step 1:

[0446] The server acquires audio data in real time from the online meeting platform. The input is the audio stream during the meeting, and the output is digital audio data stored on the server. Specifically, it uses a Web API to continuously retrieve the meeting's audio feed.

[0447] Step 2:

[0448] The server converts the acquired audio data into text data using speech recognition software. The input is the audio data acquired in step 1, and the output is character data in text format. Here, the process of analyzing the audio waveform and converting it into a string of characters in the corresponding language takes place.

[0449] Step 3:

[0450] The server analyzes the context based on the converted text data. The input is the text data generated in step 2, and the server understands its context and extracts relevant keywords. The output is a list of the analyzed contextual information and keywords. Specifically, it uses natural language processing techniques to grasp the intent and focus of the text.

[0451] Step 4:

[0452] The server uses a generative AI model based on contextual information to generate appropriate responses. The input is the contextual information analyzed in step 3, and the output is the generated response text. By providing prompt sentences to the AI ​​model, situation-appropriate questions and responses can be generated. For example, it can create responses such as, "The advantages of the new product are explained, but what other features does it have?"

[0453] Step 5:

[0454] The server sends the generated response to the terminal. The input is the response text generated in step 4, and the output is a chat-style message displayed in the user interface. The terminal visually presents the content to the user and prompts the user to take specific actions to participate in the conversation.

[0455] (Application Example 1)

[0456] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0457] In online meetings and virtual stores, there is a challenge in generating and presenting efficient and natural responses when participants or customers speak or make inquiries in real time, hindering smooth communication. Furthermore, to improve the quality of information transmission, there is a need for a system that can immediately respond to the diverse needs of participants.

[0458] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0459] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, a generation means for analyzing the context based on the converted speech text and generating an appropriate response, and an additional means for acquiring voice input from a user and generating a corresponding response based on the voice input. This makes it possible to facilitate interaction with participants and customers and improve the quality of information transmission in online meetings and virtual stores.

[0460] "Speech information" refers to audio data acquired during communication using speech, as well as information about the speaker who produced that speech.

[0461] "Acquisition means" refers to technical methods and devices for collecting voice data and necessary information from external platforms or environments.

[0462] A "conversion method" refers to a method or system for converting collected audio data into text information, which can then be used as basic data for understanding the content of a conversation.

[0463] "Generation means" refers to technologies and methods for generating natural responses and comments based on context obtained from speech-to-text.

[0464] "Display means" refers to an interface or system for visually presenting the generated response to the user.

[0465] "Additional means" refers to methods or systems that have the function of receiving direct voice input from the user, processing it, and deriving a response.

[0466] A "virtual store" is a virtual sales environment built on a digital platform as an alternative to a physical store, where users can browse products and make inquiries.

[0467] A "machine learning model" is an algorithm or network that can learn from large amounts of data and generate responses that mimic some aspects of human intelligence.

[0468] The system for realizing this invention is built on the basis of smooth data exchange between servers, terminals, and users.

[0469] The server acquires audio data from online meeting platforms and virtual store environments. Using acquisition methods, the server captures the voice spoken by the user and converts it into text using speech recognition software such as the Google Cloud Speech-to-Text API. This converted text forms the basis for contextual analysis.

[0470] Next, the server uses OpenAI's generative AI model as a generation mechanism to generate appropriate responses from the speech-to-text format. Based on prompts, it generates natural-sounding responses and questions that understand the flow and context of the conversation. For example, if there is a question about a product, the generative AI model will create a detailed response about its features and specifications.

[0471] The server then sends the generated response to the terminal for display in chat format. The terminal has a user interface through which the user can receive the response and make further inquiries. This user interface operates as an application installed on a smartphone or smart glasses and is developed using technologies such as React Native.

[0472] As a concrete example, if a user asks "Is this product waterproof?" in a virtual store, the server converts the question into text, and the generating AI model produces a response such as "Yes, this product is waterproof."

[0473] As an example of a prompt, the AI ​​model is input in the format, "Generate an appropriate response to the customer from this voice-text data: {voice-text}". Based on this prompt, a response that fits the flow of the conversation is provided, making it possible to facilitate smooth interaction with the user.

[0474] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0475] Step 1:

[0476] The server obtains speech information from the user. This information is obtained via a voice input device (e.g., a smartphone or smart glasses). The input voice data is sent to the server, and this voice data becomes the input.

[0477] Step 2:

[0478] The server converts the acquired audio data into text data using a conversion method. By utilizing the Google Cloud Speech-to-Text API, the audio is converted into text data by converting the speech into a string. The text data is then output.

[0479] Step 3:

[0480] The server uses a generative AI model to analyze the context from text data using a generation method and generate an appropriate response. At this stage, it receives text data as input and inputs prompt sentences into the generative AI model to generate a natural language response that fits the context. The response text is then generated as output.

[0481] Step 4:

[0482] The server sends the generated response to the terminal. The terminal receives this response text and displays it to the user through the user interface. A React Native application is used for display, allowing the user to see the response. A visual representation of the response is provided as output.

[0483] Step 5:

[0484] The user reviews the displayed response through the terminal's interface. They can then enter additional inquiries or comments based on that response. The user's voice or text input becomes new input, and the subsequent dialogue is processed back to step 1.

[0485] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0486] In one embodiment of the present invention, an online communication support system equipped with emotion recognition functionality is provided. This system aims to generate more empathetic and appropriate responses by processing user speech information in real time and recognizing emotions.

[0487] First, the server receives speech information from the online meeting platform. This speech information includes both audio and text data, which are acquired simultaneously. The audio data is converted to text using a speech recognition engine. The converted text is used to understand the content of the conversation.

[0488] Next, the server uses an emotion engine to analyze the user's emotions from the acquired voice and text data. This process infers the emotional state based on the tone, pitch, rhythm, and words used in the voice. For example, if the user is excited, characteristics such as a higher voice tone and faster speech will appear. This emotional information is a crucial element in making responses more human and empathetic.

[0489] Based on the emotional information obtained by the emotion engine, the server uses generation methods to produce contextually appropriate responses. The generated responses reflect the emotional state and take into account the corresponding tone and content. For example, if a user expresses anxiety, the system will generate a reassuring response such as, "Are you okay? Is there anything I can do to help?"

[0490] The generated response is sent from the server to the terminal and displayed on the user interface. The display on the terminal is designed to be intuitive and easy for the user to understand. Based on this, the user can participate in the conversation and engage in active real-time communication.

[0491] This system enables deeper mutual understanding among participants in online communication through responses that take user emotions into account. As a result, it is expected to significantly improve the efficiency and satisfaction of meetings and lectures.

[0492] The following describes the processing flow.

[0493] Step 1:

[0494] The server retrieves audio streams and metadata in real time from the online meeting platform. The audio streams, which include the speech of all participants during the meeting, are sent to the server via the platform's API.

[0495] Step 2:

[0496] The server processes the acquired audio stream through a speech recognition engine to convert it to text. This speech recognition engine uses machine learning algorithms to achieve highly accurate transcription and quickly generates text from the acquired audio data.

[0497] Step 3:

[0498] The emotion engine built into the server analyzes both audio and text data to identify the speaker's emotional state. It detects voice tone, speed, and keywords, and infers emotions (such as joy, anger, sadness, etc.) from this information. For example, a high-pitched and fast voice may indicate tension or excitement.

[0499] Step 4:

[0500] The server automatically generates an appropriate response using a generation mechanism based on the evaluation results of the emotional state and the voice-to-text. This response takes into account the user's current emotions and is designed to ensure a smooth flow of conversation. For example, if anxiety is detected, it might respond with something like, "Please calm down, and let me know if there's anything bothering you."

[0501] Step 5:

[0502] The server sends the generated response to the terminal and displays it on the user interface. The display method uses a visually simple and easy-to-understand format so that the user can quickly comprehend it.

[0503] Step 6:

[0504] Users review the responses displayed on their devices and input their own opinions and questions. Further responses and new interactions are then initiated, taking user feedback into consideration. This allows the meeting conversation to progress naturally and deepens understanding among participants.

[0505] This system facilitates smoother communication in online meetings and lectures by providing responses tailored to the user's emotional state. As a result, it is expected to improve participant satisfaction and enhance the effectiveness of the meetings.

[0506] (Example 2)

[0507] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0508] Online communication presents a challenge in instantly grasping the subtle nuances of participants' emotions and providing appropriate responses. This can lead to misunderstandings and inefficiencies in meetings and discussions. To address this, a system is needed that analyzes emotions in real time and generates empathetic responses that correspond to the user's feelings.

[0509] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0510] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into language data, and a generation means for analyzing emotions based on the converted language data and phonological features and generating a contextually appropriate response. This enables real-time, empathetic response generation that takes the user's emotions into consideration.

[0511] "Speech information" refers to conversational data during communication, including speech and associated text data.

[0512] "Acquisition means" refers to a function or method for collecting speech information from an external source and incorporating it into the system.

[0513] "Conversion means" refers to a function or method for converting acquired audio data into text data, i.e., language data.

[0514] "Language data" refers to written text information converted from audio data using speech recognition technology.

[0515] "Phonological features" refer to information that characterizes the structure of speech, including elements such as tone, pitch, and rhythm.

[0516] "Emotional analysis" is the process of using voice and language data to infer and classify the user's emotional state.

[0517] "Generating means" refers to a function or method for producing an appropriate response based on information obtained through analysis.

[0518] "Contextually appropriate responses" refer to natural and empathetic responses that are based on the user's statements and emotional state.

[0519] "Display means" refers to a function or method for visually presenting the generated response on a user interface.

[0520] This invention relates to an online communication support system equipped with emotion recognition capabilities. The system aims to collect user speech information in real time, analyze emotions, and generate appropriate responses.

[0521] The server acquires speech information from the online meeting platform. This information includes audio and text data, which are incorporated into the system through the acquisition mechanism. Using speech recognition software (e.g., a speech recognition API), the server converts the audio data into text data. The converted text data serves as the basis for analyzing the content of the conversation.

[0522] Next, the server uses emotion pattern recognition software (e.g., emotion analysis engine) to analyze the phonological features contained in the audio data. This analysis infers and classifies the speaker's emotional state. Based on this information, the server utilizes a generative AI model (e.g., natural language generation model) to generate a contextually appropriate response.

[0523] The generated response is sent to the terminal via the network. The terminal displays the response in its user interface, allowing the user to understand it immediately and facilitate smooth communication.

[0524] For example, if a user says, "I'm worried because the project is behind schedule," the server performs sentiment analysis on that utterance and assigns a label indicating anxiety. Next, it uses a generative AI model to generate a reassuring response such as, "Let's think together about how we can support you."

[0525] Examples of prompts to input into a generative AI model:

[0526] User utterance: "I'm worried because the project is behind schedule."

[0527] Emotional labeling based on analysis: Anxiety

[0528] Response generation prompt: "What kind of reassurance can we provide to anxious users?"

[0529] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0530] Step 1:

[0531] The server acquires speech information in real time from the online meeting platform. The input consists of audio data and text data, and it receives both types of data and begins processing them simultaneously. The audio data is sent to the speech recognition engine for processing, and the text data is sent to the sentiment analysis engine.

[0532] Step 2:

[0533] The server uses speech recognition software to convert the received audio data into language data. Specifically, it converts the audio file into a byte stream format and sends it to the speech recognition API. The input is audio data, and the output is string data that reflects the content of the audio. This text is used in subsequent parsing steps.

[0534] Step 3:

[0535] The server analyzes emotions using an emotion analysis engine based on linguistic data and phonological features. It analyzes voice tone, pitch, and rhythm with high accuracy and performs semantic analysis on this data along with text data. The input consists of phonological feature parameters and text data, and the output consists of emotion labels (e.g., joy, sadness, anger) and numerical data indicating their intensity.

[0536] Step 4:

[0537] Using a generative AI model, appropriate responses are generated using analyzed sentiment labels. The server prepares sentiment labels and text data as prompts and inputs them into the AI ​​model. Based on this input, the AI ​​model generates a contextually appropriate and natural response. The input consists of sentiment labels and prompts, and the output is text data to be sent back to the user.

[0538] Step 5:

[0539] The server sends the generated response to the terminal. The terminal displays the received response text in its user interface. Specifically, it creates a notification on the screen based on the response text and presents it to the user in an easily understandable format, either audibly or visually. The input is the generated response text, and the output is the visible display shown on the screen.

[0540] (Application Example 2)

[0541] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0542] In autonomous vehicles, there is a challenge in providing an appropriate in-vehicle environment that responds to the emotional state of passengers. If it were possible to respond to passengers' emotions, travel in autonomous vehicles would be more comfortable and safer. Conventional technologies have been unable to accurately grasp passengers' emotions and provide responses based on them in real time.

[0543] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0544] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, an analysis means for analyzing emotions based on the converted speech text and speech features, a generation means for generating an appropriate response based on the emotion information obtained by the analysis, and a display means for displaying the generated response on a user input device. This enables real-time responses and adjustments to the in-vehicle environment in accordance with the emotional state of passengers.

[0545] "Speech information" refers to audio waves and associated text data, which are data acquired during communication.

[0546] "Acquisition means" refers to a device or method for receiving underlying data from voice input devices or sensors.

[0547] "Speech-to-text" refers to textual information generated from sound waves using speech recognition technology.

[0548] "Conversion means" refers to technology or equipment for converting sound waves into textual information.

[0549] "Means of analyzing emotions" refers to technologies or devices for inferring the emotional state of a speaker from voice and text data.

[0550] "Generation means" refers to a method or apparatus for creating appropriate responses or actions based on analyzed data.

[0551] "Display means" refers to a method or apparatus for providing a generated response to a user visually or audibly.

[0552] A "user input device" refers to a device used by a user to receive or manipulate information.

[0553] The system for carrying out this invention comprises several technical elements. The server collects sound wave data from microphones installed in the vehicle using an acquisition means for acquiring passenger speech information. This sound wave data is then converted into speech-to-text by a conversion means and processed as text data.

[0554] The server uses analytical tools responsible for sentiment analysis to evaluate the passenger's emotional state from the characteristics of the speech text and sound waves. This analysis takes into account factors such as the tone, tempo, and intensity of the speech, as well as any bias in the vocabulary used. The results of the analysis are recorded in the server as the passenger's emotional state.

[0555] Next, the server uses a generation mechanism to generate a response based on the acquired emotional information. This response includes in-car voice announcements that take the passenger's emotions into consideration, automatically selected background music, and voice guidance. For example, if it is determined that a passenger is feeling tense, relaxing background music may be automatically played.

[0556] The responses displayed on the terminal are provided by a display mechanism and presented in a format that is easy for passengers to understand and use. Through this series of technological elements, passenger emotional satisfaction is enhanced, enabling a safe and comfortable autonomous vehicle experience.

[0557] For example, if the conversation in the car during a family trip is cheerful and positive, the system will select upbeat background music and take measures to enhance the atmosphere inside the car. Another example of a prompt to be input to the generation AI model is in the format of "Detect the emotion of the following sentence and generate the corresponding feedback: Text." This allows the system to identify the passenger's emotion from the text and provide the most appropriate in-car response.

[0558] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0559] Step 1:

[0560] The server acquires passenger speech information from microphones installed inside the vehicle. The input is audio wave data. The server collects this audio wave data using an acquisition method and stores it in a format that can be processed in the next step.

[0561] Step 2:

[0562] The server converts audio wave data into speech-to-text using a conversion mechanism. The input is audio wave data, and the output is speech-to-text. In this process, speech recognition software analyzes the characteristics of the sound waves and generates corresponding character information. Specifically, it analyzes the amplitude and frequency of the sound waves and converts them into Japanese words and phrases.

[0563] Step 3:

[0564] The server analyzes the characteristics of the generated speech-text and speech-wave data and uses analytical means to estimate the passenger's emotions. The input is speech-text and speech-wave features, and the output is emotion data. This emotion analysis takes into account voice tone, speed, and vocabulary used to estimate a specific emotional state (e.g., joy, sadness, tension).

[0565] Step 4:

[0566] The server utilizes generation methods to generate appropriate responses based on the passenger's emotional state, using emotion analysis results. The input is emotion data, and the output is response text and associated in-vehicle operation commands. This step generates text for emotion-appropriate background music playback and voice guidance. A generation AI model is used to design natural, situation-aware responses.

[0567] Step 5:

[0568] The server sends the generated response to the terminal and displays it on the in-vehicle user interface. Inputs are response text and in-vehicle operation commands, and outputs are the in-vehicle responses perceived by the user (music playback, voice guidance, etc.). The terminal provides passengers with specific responses via voice, enabling them to take action or make choices based on them.

[0569] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0570] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0571] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0572] [Fourth Embodiment]

[0573] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0574] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0575] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0576] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0577] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0578] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0579] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0580] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0581] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0582] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0583] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0584] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0585] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0586] In one embodiment of the invention, a system is constructed to improve real-time communication in online meetings and lectures. The specific operation of this system is described below.

[0587] First, the server acquires speech information in real time from the online meeting platform. This information includes voice data and speaker IDs, and is continuously collected as the meeting progresses.

[0588] Next, the server converts the acquired audio data into text format using speech recognition technology. This converted audio-text becomes the basis for analyzing the content of the conversation. For example, if the speaker is discussing a new product, relevant keywords and sentences will be reflected in the text.

[0589] Next, the server's generation system analyzes the context using the generated text and produces appropriate responses and comments. This generation process utilizes machine learning models to generate natural interjections and questions that fit the flow of the conversation. For example, if the speaker is talking about a new feature, the AI ​​might create a response such as, "How will that feature improve our work?"

[0590] The generated response is immediately sent from the server to the terminal. On the terminal, this response is displayed in a chat format via a user interface, allowing participants to review it and enter their own opinions or questions as needed. This user interface is designed with ease of use and visibility in mind, making it intuitive for participants to use.

[0591] Furthermore, by allowing users to participate in the discussion using interjections and comments provided through this system, speakers can instantly grasp the audience's interest and level of understanding. This enables speakers to adjust the information they provide in real time, which is expected to improve the quality of online meetings.

[0592] By utilizing such a system, smooth communication can be promoted for both speakers and listeners, enabling effective information transmission and exchange of opinions even in an online environment.

[0593] The following describes the processing flow.

[0594] Step 1:

[0595] The server retrieves audio streams in real time from the online meeting platform. This includes connecting to the platform via an API and receiving audio data. The audio data is accompanied by metadata such as speaker information.

[0596] Step 2:

[0597] The server inputs the acquired audio stream into a speech recognition engine, which converts it into text format. The speech recognition engine utilizes machine learning models to perform highly accurate text conversion. At this stage, the meeting content is saved as text.

[0598] Step 3:

[0599] The server analyzes the speech-to-text data to understand the context of the conversation. Natural language processing algorithms extract keywords and topics to help understand the content of the conversation. This analysis is then used to generate subsequent responses.

[0600] Step 4:

[0601] The server uses a generation mechanism to generate appropriate responses and interjections based on the analyzed context. This process utilizes a generative AI model, which is required to create responses that naturally match the context and flow of the conversation. For example, responses such as "That's interesting, could you tell me more?" might be generated.

[0602] Step 5:

[0603] The server sends the generated response to the terminal. Since transmission occurs in real time, the communication protocol is optimized to minimize delays.

[0604] Step 6:

[0605] The terminal displays received responses in the user interface. The user interface is designed to allow users to view responses in the form of a chat window or similar. Participants can view these responses and interact as needed.

[0606] Step 7:

[0607] Users offer their opinions and ask questions based on the displayed responses. This interaction makes the conversation more lively and facilitates smoother communication with the speaker. User input is also collected as a log and used to improve the system.

[0608] (Example 1)

[0609] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0610] Achieving smooth, real-time communication in online meetings remains challenging. In voice-based conversations, it's difficult for all participants to understand and react simultaneously, often resulting in one-way dialogue. This can hinder the stimulating nature of discussions and the efficient transmission of information. Therefore, there is a need for a system that instantly transcribes spoken content into text and generates appropriate responses on the spot.

[0611] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0612] In this invention, the server includes an acquisition means for collecting audio data via a communication means linked to an electronic device, a conversion means for converting the acquired audio data into text data, and a generation means for understanding the context based on the converted text data and generating an appropriate response. This makes it possible to transcribe speech during a meeting in real time and quickly provide natural responses appropriate to the context.

[0613] "Electronic devices" is a general term for devices used to process or transmit digital information.

[0614] "Communication methods" refer to methods and protocols for sending and receiving information, and include mechanisms for exchanging data over a network.

[0615] "Audio data" refers to information that is recorded or transmitted using digital methods, specifically human speech.

[0616] "Means of acquisition" refers to the process or method for collecting specific information.

[0617] "Text data" refers to information expressed in text format, including audio transcribed into text.

[0618] "Conversion means" refers to technologies and processes for changing data formats or the nature of information.

[0619] "Context" refers to the situation and relationships necessary to understand the meaning and information behind a sentence or statement.

[0620] "Generative means" refers to the technologies and processes used to create new information and data.

[0621] An "information display device" refers to hardware or an interface used to visually present information content such as text and graphics.

[0622] This invention is a system for streamlining real-time communication in online environments. Specifically, it facilitates communication by instantly converting spoken information into text and generating responses during online meetings.

[0623] The server retrieves audio data from the electronic conferencing platform. A common API can be used for this task. For example, when using a web conferencing platform, audio data can be collected in real time through its API.

[0624] The acquired audio data is converted into text data on the server using speech recognition software, such as the Google Speech-to-Text API. This data forms the basis for subsequent analysis and response generation.

[0625] The server utilizes the converted text data to analyze the context using a generative AI model, such as a machine learning algorithm provided by a business partner, and generates an appropriate response. This model is customized to obtain the necessary response using prompt sentences. For example, it might use a prompt sentence like, "Generate appropriate questions for when the speaker is explaining a new product."

[0626] The generated responses are sent to the terminal and displayed on the user interface via an information display device. This interface is designed to be intuitive for users, allowing participants to express their opinions in real time based on the generated responses.

[0627] This allows users to communicate actively even in online meeting environments, thereby improving the quality of meetings.

[0628] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0629] Step 1:

[0630] The server acquires audio data in real time from the online meeting platform. The input is the audio stream during the meeting, and the output is digital audio data stored on the server. Specifically, it uses a Web API to continuously retrieve the meeting's audio feed.

[0631] Step 2:

[0632] The server converts the acquired audio data into text data using speech recognition software. The input is the audio data acquired in step 1, and the output is character data in text format. Here, the process of analyzing the audio waveform and converting it into a string of characters in the corresponding language takes place.

[0633] Step 3:

[0634] The server analyzes the context based on the converted text data. The input is the text data generated in step 2, and the server understands its context and extracts relevant keywords. The output is a list of the analyzed contextual information and keywords. Specifically, it uses natural language processing techniques to grasp the intent and focus of the text.

[0635] Step 4:

[0636] The server uses a generative AI model based on contextual information to generate appropriate responses. The input is the contextual information analyzed in step 3, and the output is the generated response text. By providing prompt sentences to the AI ​​model, situation-appropriate questions and responses can be generated. For example, it can create responses such as, "The advantages of the new product are explained, but what other features does it have?"

[0637] Step 5:

[0638] The server sends the generated response to the terminal. The input is the response text generated in step 4, and the output is a chat-style message displayed in the user interface. The terminal visually presents the content to the user and prompts the user to take specific actions to participate in the conversation.

[0639] (Application Example 1)

[0640] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0641] In online meetings and virtual stores, there is a challenge in generating and presenting efficient and natural responses when participants or customers speak or make inquiries in real time, hindering smooth communication. Furthermore, to improve the quality of information transmission, there is a need for a system that can immediately respond to the diverse needs of participants.

[0642] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0643] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, a generation means for analyzing the context based on the converted speech text and generating an appropriate response, and an additional means for acquiring voice input from a user and generating a corresponding response based on the voice input. This makes it possible to facilitate interaction with participants and customers and improve the quality of information transmission in online meetings and virtual stores.

[0644] "Speech information" refers to audio data acquired during communication using speech, as well as information about the speaker who produced that speech.

[0645] "Acquisition means" refers to technical methods and devices for collecting voice data and necessary information from external platforms or environments.

[0646] A "conversion method" refers to a method or system for converting collected audio data into text information, which can then be used as basic data for understanding the content of a conversation.

[0647] "Generation means" refers to technologies and methods for generating natural responses and comments based on context obtained from speech-to-text.

[0648] "Display means" refers to an interface or system for visually presenting the generated response to the user.

[0649] "Additional means" refers to methods or systems that have the function of receiving direct voice input from the user, processing it, and deriving a response.

[0650] A "virtual store" is a virtual sales environment built on a digital platform as an alternative to a physical store, where users can browse products and make inquiries.

[0651] A "machine learning model" is an algorithm or network that can learn from large amounts of data and generate responses that mimic some aspects of human intelligence.

[0652] The system for realizing this invention is built on the basis of smooth data exchange between servers, terminals, and users.

[0653] The server acquires audio data from online meeting platforms and virtual store environments. Using acquisition methods, the server captures the voice spoken by the user and converts it into text using speech recognition software such as the Google Cloud Speech-to-Text API. This converted text forms the basis for contextual analysis.

[0654] Next, the server uses OpenAI's generative AI model as a generation mechanism to generate appropriate responses from the speech-to-text format. Based on prompts, it generates natural-sounding responses and questions that understand the flow and context of the conversation. For example, if there is a question about a product, the generative AI model will create a detailed response about its features and specifications.

[0655] The server then sends the generated response to the terminal for display in chat format. The terminal has a user interface through which the user can receive the response and make further inquiries. This user interface operates as an application installed on a smartphone or smart glasses and is developed using technologies such as React Native.

[0656] As a concrete example, if a user asks "Is this product waterproof?" in a virtual store, the server converts the question into text, and the generating AI model produces a response such as "Yes, this product is waterproof."

[0657] As an example of a prompt, the AI ​​model is input in the format, "Generate an appropriate response to the customer from this voice-text data: {voice-text}". Based on this prompt, a response that fits the flow of the conversation is provided, making it possible to facilitate smooth interaction with the user.

[0658] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0659] Step 1:

[0660] The server obtains speech information from the user. This information is obtained via a voice input device (e.g., a smartphone or smart glasses). The input voice data is sent to the server, and this voice data becomes the input.

[0661] Step 2:

[0662] The server converts the acquired audio data into text data using a conversion method. By utilizing the Google Cloud Speech-to-Text API, the audio is converted into text data by converting the speech into a string. The text data is then output.

[0663] Step 3:

[0664] The server uses a generative AI model to analyze the context from text data using a generation method and generate an appropriate response. At this stage, it receives text data as input and inputs prompt sentences into the generative AI model to generate a natural language response that fits the context. The response text is then generated as output.

[0665] Step 4:

[0666] The server sends the generated response to the terminal. The terminal receives this response text and displays it to the user through the user interface. A React Native application is used for display, allowing the user to see the response. A visual representation of the response is provided as output.

[0667] Step 5:

[0668] The user reviews the displayed response through the terminal's interface. They can then enter additional inquiries or comments based on that response. The user's voice or text input becomes new input, and the subsequent dialogue is processed back to step 1.

[0669] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0670] In one embodiment of the present invention, an online communication support system equipped with emotion recognition functionality is provided. This system aims to generate more empathetic and appropriate responses by processing user speech information in real time and recognizing emotions.

[0671] First, the server receives speech information from the online meeting platform. This speech information includes both audio and text data, which are acquired simultaneously. The audio data is converted to text using a speech recognition engine. The converted text is used to understand the content of the conversation.

[0672] Next, the server uses an emotion engine to analyze the user's emotions from the acquired voice and text data. This process infers the emotional state based on the tone, pitch, rhythm, and words used in the voice. For example, if the user is excited, characteristics such as a higher voice tone and faster speech will appear. This emotional information is a crucial element in making responses more human and empathetic.

[0673] Based on the emotional information obtained by the emotion engine, the server uses generation methods to produce contextually appropriate responses. The generated responses reflect the emotional state and take into account the corresponding tone and content. For example, if a user expresses anxiety, the system will generate a reassuring response such as, "Are you okay? Is there anything I can do to help?"

[0674] The generated response is sent from the server to the terminal and displayed on the user interface. The display on the terminal is designed to be intuitive and easy for the user to understand. Based on this, the user can participate in the conversation and engage in active real-time communication.

[0675] This system enables deeper mutual understanding among participants in online communication through responses that take user emotions into account. As a result, it is expected to significantly improve the efficiency and satisfaction of meetings and lectures.

[0676] The following describes the processing flow.

[0677] Step 1:

[0678] The server retrieves audio streams and metadata in real time from the online meeting platform. The audio streams, which include the speech of all participants during the meeting, are sent to the server via the platform's API.

[0679] Step 2:

[0680] The server processes the acquired audio stream through a speech recognition engine to convert it to text. This speech recognition engine uses machine learning algorithms to achieve highly accurate transcription and quickly generates text from the acquired audio data.

[0681] Step 3:

[0682] The emotion engine built into the server analyzes both audio and text data to identify the speaker's emotional state. It detects voice tone, speed, and keywords, and infers emotions (such as joy, anger, sadness, etc.) from this information. For example, a high-pitched and fast voice may indicate tension or excitement.

[0683] Step 4:

[0684] The server automatically generates an appropriate response using a generation mechanism based on the evaluation results of the emotional state and the voice-to-text. This response takes into account the user's current emotions and is designed to ensure a smooth flow of conversation. For example, if anxiety is detected, it might respond with something like, "Please calm down, and let me know if there's anything bothering you."

[0685] Step 5:

[0686] The server sends the generated response to the terminal and displays it on the user interface. The display method uses a visually simple and easy-to-understand format so that the user can quickly comprehend it.

[0687] Step 6:

[0688] Users review the responses displayed on their devices and input their own opinions and questions. Further responses and new interactions are then initiated, taking user feedback into consideration. This allows the meeting conversation to progress naturally and deepens understanding among participants.

[0689] This system facilitates smoother communication in online meetings and lectures by providing responses tailored to the user's emotional state. As a result, it is expected to improve participant satisfaction and enhance the effectiveness of the meetings.

[0690] (Example 2)

[0691] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0692] Online communication presents a challenge in instantly grasping the subtle nuances of participants' emotions and providing appropriate responses. This can lead to misunderstandings and inefficiencies in meetings and discussions. To address this, a system is needed that analyzes emotions in real time and generates empathetic responses that correspond to the user's feelings.

[0693] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0694] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into language data, and a generation means for analyzing emotions based on the converted language data and phonological features and generating a contextually appropriate response. This enables real-time, empathetic response generation that takes the user's emotions into consideration.

[0695] "Speech information" refers to conversational data during communication, including speech and associated text data.

[0696] "Acquisition means" refers to a function or method for collecting speech information from an external source and incorporating it into the system.

[0697] "Conversion means" refers to a function or method for converting acquired audio data into text data, i.e., language data.

[0698] "Language data" refers to written text information converted from audio data using speech recognition technology.

[0699] "Phonological features" refer to information that characterizes the structure of speech, including elements such as tone, pitch, and rhythm.

[0700] "Emotional analysis" is the process of using voice and language data to infer and classify the user's emotional state.

[0701] "Generating means" refers to a function or method for producing an appropriate response based on information obtained through analysis.

[0702] "Contextually appropriate responses" refer to natural and empathetic responses that are based on the user's statements and emotional state.

[0703] "Display means" refers to a function or method for visually presenting the generated response on a user interface.

[0704] This invention relates to an online communication support system equipped with emotion recognition capabilities. The system aims to collect user speech information in real time, analyze emotions, and generate appropriate responses.

[0705] The server acquires speech information from the online meeting platform. This information includes audio and text data, which are incorporated into the system through the acquisition mechanism. Using speech recognition software (e.g., a speech recognition API), the server converts the audio data into text data. The converted text data serves as the basis for analyzing the content of the conversation.

[0706] Next, the server uses emotion pattern recognition software (e.g., emotion analysis engine) to analyze the phonological features contained in the audio data. This analysis infers and classifies the speaker's emotional state. Based on this information, the server utilizes a generative AI model (e.g., natural language generation model) to generate a contextually appropriate response.

[0707] The generated response is sent to the terminal via the network. The terminal displays the response in its user interface, allowing the user to understand it immediately and facilitate smooth communication.

[0708] For example, if a user says, "I'm worried because the project is behind schedule," the server performs sentiment analysis on that utterance and assigns a label indicating anxiety. Next, it uses a generative AI model to generate a reassuring response such as, "Let's think together about how we can support you."

[0709] Examples of prompts to input into a generative AI model:

[0710] User utterance: "I'm worried because the project is behind schedule."

[0711] Emotional labeling based on analysis: Anxiety

[0712] Response generation prompt: "What kind of reassurance can we provide to anxious users?"

[0713] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0714] Step 1:

[0715] The server acquires speech information in real time from the online meeting platform. The input consists of audio data and text data, and it receives both types of data and begins processing them simultaneously. The audio data is sent to the speech recognition engine for processing, and the text data is sent to the sentiment analysis engine.

[0716] Step 2:

[0717] The server uses speech recognition software to convert the received audio data into language data. Specifically, it converts the audio file into a byte stream format and sends it to the speech recognition API. The input is audio data, and the output is string data that reflects the content of the audio. This text is used in subsequent parsing steps.

[0718] Step 3:

[0719] The server analyzes emotions using an emotion analysis engine based on linguistic data and phonological features. It analyzes voice tone, pitch, and rhythm with high accuracy and performs semantic analysis on this data along with text data. The input consists of phonological feature parameters and text data, and the output consists of emotion labels (e.g., joy, sadness, anger) and numerical data indicating their intensity.

[0720] Step 4:

[0721] Using a generative AI model, appropriate responses are generated using analyzed sentiment labels. The server prepares sentiment labels and text data as prompts and inputs them into the AI ​​model. Based on this input, the AI ​​model generates a contextually appropriate and natural response. The input consists of sentiment labels and prompts, and the output is text data to be sent back to the user.

[0722] Step 5:

[0723] The server sends the generated response to the terminal. The terminal displays the received response text in its user interface. Specifically, it creates a notification on the screen based on the response text and presents it to the user in an easily understandable format, either audibly or visually. The input is the generated response text, and the output is the visible display shown on the screen.

[0724] (Application Example 2)

[0725] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0726] In autonomous vehicles, there is a challenge in providing an appropriate in-vehicle environment that responds to the emotional state of passengers. If it were possible to respond to passengers' emotions, travel in autonomous vehicles would be more comfortable and safer. Conventional technologies have been unable to accurately grasp passengers' emotions and provide responses based on them in real time.

[0727] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0728] In this invention, the server includes an acquisition means for acquiring speech information, a conversion means for converting the acquired speech information into speech text, an analysis means for analyzing emotions based on the converted speech text and speech features, a generation means for generating an appropriate response based on the emotion information obtained by the analysis, and a display means for displaying the generated response on a user input device. This enables real-time responses and adjustments to the in-vehicle environment in accordance with the emotional state of passengers.

[0729] "Speech information" refers to audio waves and associated text data, which are data acquired during communication.

[0730] "Acquisition means" refers to a device or method for receiving underlying data from voice input devices or sensors.

[0731] "Speech-to-text" refers to textual information generated from sound waves using speech recognition technology.

[0732] "Conversion means" refers to technology or equipment for converting sound waves into textual information.

[0733] "Means of analyzing emotions" refers to technologies or devices for inferring the emotional state of a speaker from voice and text data.

[0734] "Generation means" refers to a method or apparatus for creating appropriate responses or actions based on analyzed data.

[0735] "Display means" refers to a method or apparatus for providing a generated response to a user visually or audibly.

[0736] A "user input device" refers to a device used by a user to receive or manipulate information.

[0737] The system for carrying out this invention comprises several technical elements. The server collects sound wave data from microphones installed in the vehicle using an acquisition means for acquiring passenger speech information. This sound wave data is then converted into speech-to-text by a conversion means and processed as text data.

[0738] The server uses analytical tools responsible for sentiment analysis to evaluate the passenger's emotional state from the characteristics of the speech text and sound waves. This analysis takes into account factors such as the tone, tempo, and intensity of the speech, as well as any bias in the vocabulary used. The results of the analysis are recorded in the server as the passenger's emotional state.

[0739] Next, the server uses a generation mechanism to generate a response based on the acquired emotional information. This response includes in-car voice announcements that take the passenger's emotions into consideration, automatically selected background music, and voice guidance. For example, if it is determined that a passenger is feeling tense, relaxing background music may be automatically played.

[0740] The responses displayed on the terminal are provided by a display mechanism and presented in a format that is easy for passengers to understand and use. Through this series of technological elements, passenger emotional satisfaction is enhanced, enabling a safe and comfortable autonomous vehicle experience.

[0741] For example, if the conversation in the car during a family trip is cheerful and positive, the system will select upbeat background music and take measures to enhance the atmosphere inside the car. Another example of a prompt to be input to the generation AI model is in the format of "Detect the emotion of the following sentence and generate the corresponding feedback: Text." This allows the system to identify the passenger's emotion from the text and provide the most appropriate in-car response.

[0742] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0743] Step 1:

[0744] The server acquires passenger speech information from microphones installed inside the vehicle. The input is audio wave data. The server collects this audio wave data using an acquisition method and stores it in a format that can be processed in the next step.

[0745] Step 2:

[0746] The server converts audio wave data into speech-to-text using a conversion mechanism. The input is audio wave data, and the output is speech-to-text. In this process, speech recognition software analyzes the characteristics of the sound waves and generates corresponding character information. Specifically, it analyzes the amplitude and frequency of the sound waves and converts them into Japanese words and phrases.

[0747] Step 3:

[0748] The server analyzes the characteristics of the generated speech-text and speech-wave data and uses analytical means to estimate the passenger's emotions. The input is speech-text and speech-wave features, and the output is emotion data. This emotion analysis takes into account voice tone, speed, and vocabulary used to estimate a specific emotional state (e.g., joy, sadness, tension).

[0749] Step 4:

[0750] The server utilizes generation methods to generate appropriate responses based on the passenger's emotional state, using emotion analysis results. The input is emotion data, and the output is response text and associated in-vehicle operation commands. This step generates text for emotion-appropriate background music playback and voice guidance. A generation AI model is used to design natural, situation-aware responses.

[0751] Step 5:

[0752] The server sends the generated response to the terminal and displays it on the in-vehicle user interface. Inputs are response text and in-vehicle operation commands, and outputs are the in-vehicle responses perceived by the user (music playback, voice guidance, etc.). The terminal provides passengers with specific responses via voice, enabling them to take action or make choices based on them.

[0753] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0754] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0755] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0756] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0757] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0758] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0759] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0760] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0761] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0762] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0763] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0764] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0765] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0766] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0767] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0768] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0769] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0770] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0771] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0772] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0773] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0774] The following is further disclosed regarding the embodiments described above.

[0775] (Claim 1)

[0776] A means for acquiring speech information,

[0777] A conversion means for converting the acquired speech information into speech text,

[0778] A generation means that analyzes the context based on the converted speech text and generates an appropriate response,

[0779] A display means for displaying the generated response on the user interface,

[0780] A system that includes this.

[0781] (Claim 2)

[0782] The system according to claim 1, wherein the acquisition means acquires speech information from an online meeting platform.

[0783] (Claim 3)

[0784] The system according to claim 1, wherein the generation means generates a response using a machine learning model.

[0785] "Example 1"

[0786] (Claim 1)

[0787] An acquisition means for collecting voice data via a communication means linked to an electronic device,

[0788] A conversion means for converting the acquired audio data into text data,

[0789] A generation means that grasps the context based on the converted text data and generates an appropriate response,

[0790] A display means for visualizing the generated reaction on an information display device,

[0791] A system that includes this.

[0792] (Claim 2)

[0793] The system according to claim 1, wherein the acquisition means collects audio data from an electronic conferencing platform.

[0794] (Claim 3)

[0795] The system according to claim 1, wherein the generation means generates a response using a learning algorithm.

[0796] "Application Example 1"

[0797] (Claim 1)

[0798] A means for acquiring speech information,

[0799] A conversion means for converting the acquired speech information into speech text,

[0800] A generation means that analyzes the context based on the converted speech text and generates an appropriate response,

[0801] A display means for displaying the generated response on the user interface,

[0802] An additional means for acquiring voice input from a user and generating a corresponding response based on said voice input,

[0803] A means for generating and presenting appropriate responses to customer inquiries in a virtual store,

[0804] A system that includes this.

[0805] (Claim 2)

[0806] The system according to claim 1, wherein the acquisition means acquires speech information from an online meeting platform or a virtual store environment.

[0807] (Claim 3)

[0808] The system according to claim 1, wherein the generation means uses a machine learning model to generate a response based on the context of the conversation and provides it in a format suitable for use in a virtual store.

[0809] "Example 2 of combining an emotion engine"

[0810] (Claim 1)

[0811] A means for acquiring speech information,

[0812] A conversion means for converting the acquired speech information into language data,

[0813] A generation means that analyzes emotions based on the converted language data and phonological features and generates a contextually appropriate response,

[0814] A display means that transmits the generated response to the user interface and displays it,

[0815] A system that includes this.

[0816] (Claim 2)

[0817] The system according to claim 1, wherein the acquisition means acquires speech information from a communication conferencing infrastructure.

[0818] (Claim 3)

[0819] The system according to claim 1, wherein the generation means uses a learning model to generate a response based on the analysis results.

[0820] "Application example 2 when combining with an emotional engine"

[0821] (Claim 1)

[0822] A means for acquiring speech information,

[0823] A conversion means for converting the acquired speech information into speech text,

[0824] An analysis means for analyzing emotions based on the converted speech text and speech characteristics,

[0825] A generation means for generating an appropriate response based on the emotional information obtained by the analysis,

[0826] A display means for displaying the generated response on a user input device,

[0827] A system that includes this.

[0828] (Claim 2)

[0829] The system according to claim 1, wherein the acquisition means acquires speech information from a voice input device in a temporary transport means.

[0830] (Claim 3)

[0831] The system according to claim 1, wherein the generating means generates a response using a machine model. [Explanation of Symbols]

[0832] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring speech information, A conversion means for converting the acquired speech information into speech text, A generation means that analyzes the context based on the converted speech text and generates an appropriate response, A display means for displaying the generated response on the user interface, A system that includes this.

2. The system according to claim 1, wherein the acquisition means acquires speech information from an online meeting platform.

3. The system according to claim 1, wherein the generation means generates a response using a machine learning model.